You Just Hired a Million "Bad Employees"

For the first time in history, humans are cheaper than software.

01. After "Humans Are Cheaper Than Software"

For the first time in history, humans are cheaper than software.

This line comes from a long post Hebbia CEO George Sivulka published last week. Marc Andreessen retweeted it the same day with a single word — "Interesting" — and the original post has since racked up over 2.15 million views.

Much of the article repurposes a vocabulary once used to diagnose corporate management dysfunction, reapplying it to name the problems of the agent era:

  • Token-maxxing is the new "throwing bodies at the problem"
  • Loop is the new "meeting that could have been an email"
  • Wasted tokens are AI's version of organizational bloat
  • Eval (an automated scoring system for AI outputs) is the new OKR...

The post lands at the intersection of two debates that have consumed Silicon Valley over the past quarter. The first half was token-maxxing: employees at Meta and other major tech companies competed on internal leaderboards to see who could burn through the most tokens. The second half was the backlash: Gary Marcus cited data showing H200 rental prices had dropped 40% in three weeks; Fortune declared token-maxxing dead...

Every few months, AI circles also spawn some new "XX engineering." First prompt engineering; then context engineering; then harness engineering; now loop engineering... Sivulka takes a different angle: these are fundamentally management problems, not just engineering problems.

He opens with a train collision in 1841. U.S. railway mileage had exploded 120-fold in a decade, but scheduling still relied entirely on individual conductors' judgment — a complexity far beyond human capacity, culminating in a fatal crash. Railroad companies were forced to invent an entire system of management. Modern management was born, and railways became an industry commanding 60% of total stock market value.

History is repeating itself. Today we've given every employee, including the worst performer, the power to summon unlimited AI agents and burn through unlimited compute budgets.

The result? He estimates 80% of tokens are wasted, because we still haven't defined what "good" looks like.

His proposed solution is eval: building automated scoring systems for every business scenario, so AI output has standards to meet — just as code either runs or throws an error.

The essay ultimately lands on an investment thesis: the next trillion-dollar opportunity won't belong to "neofirms" trying to eat the $21 trillion services economy, but to "AI transformation companies" that help incumbent enterprises encode their implicit processes into evals and agents.

He points to Palantir as the template. At a market cap of over $300 billion, Palantir should have been the first to zero out if the "AI makes SaaS worthless" disruption thesis held — Claude can write applications now, after all. But it hasn't, because Palantir never sold software; it sold transformation: embedding inside a company, mastering its business, encoding its processes into systems. That's what the next trillion-dollar opportunity looks like.

In the grand debate over AI and employment, Jeff Bezos says AI will create labor shortages; Geoffrey Hinton warns it will eliminate vast numbers of jobs. Sivulka sides with Bezos, adding one condition: new jobs will come, but only if someone learns to manage.

As for how to manage, Sivulka frames his answer as seven parallels, with one core insight at the center: agent teams and human teams fail in exactly the same ways.


02. Full Translation: You Just Hired a Million Bad Employees

By George Sivulka (Founder & CEO, Hebbia)

@gsivulka

Everyone thought AI would replace human labor. The opposite happened.

For the first time in history, humans are cheaper than software (because token costs are staggering).

Per-capita token spend at top companies

And AI has created more jobs than it destroyed.

Employee headcount growth after AI adoption

Every time technology solves one problem, it creates another.

In the 1830s, railroads unleashed an infrastructure wave the world had never seen. American railway mileage surged 120-fold in a decade.

Then the system broke.

On October 5, 1841, on the Western Railroad in Massachusetts, two trains collided head-on due to a simple scheduling error — a fatal disaster.

As the rail network grew more complex, individual conductors could no longer ensure safety. So railroad companies embarked on decades of effort: assigning managers to each region, creating new roles within the organization, building hierarchical structures with clear reporting lines.

Modern management was born. With this system, railways became the world's first billion-dollar industry, peaking at roughly 60% of total stock market value.

Now AI is breaking the system again.

We just gave every employee, even the worst one, unlimited "headcount" and unlimited budget.

Managing AI is harder than managing people, because AI scales every failure mode instantly. Fortunately, history offers a precedent:

Agent teams and human teams fail in exactly the same ways.

Understanding these seven key parallels will unlock the next trillion dollars of AI value creation.

I. Token-maxxing is the new "throwing bodies at the problem"

Token-maxxing went from viral sensation to punchline in less than a month.

But the problem was never how many tokens got burned.

People spend so much on tokens because they have no idea what else to do.

Maybe one in a hundred employees actually understands how to feed context to AI. These people are a rare breed: they can articulate processes clearly, and have the patience to accommodate contaminated context windows — and frankly, not many even understand what "contaminated context window" means.

Hand an agent harness to the other ninety-nine, and all they'll produce is "loop."

II. Loop is the new "meeting that could have been an email"

Whether in Claude Code/Cowork, Copilot, Karpathy's Autoresearch, or any other harness, loop is just a band-aid. It conceals the fact that almost nobody can write a decent prompt.

Loop is brute force compensating for human inadequacy. The reason an agent has to call itself repeatedly to correct itself is that the human never clearly defined the task in the first place — so brute force becomes the only path forward. And the root of it all is that the human never truly understood the task to begin with.

You're burning tokens for the sake of burning tokens.

III. Wasted tokens are AI's version of organizational bloat

Most companies today are poorly managed.

Many employees have no material impact on the business. They're gears in a machine: stamping approval at every layer, then hiring more gears to keep a machine running for its own sake.

They're just "looping."

Often, cutting the loop entirely is more efficient. Musk fired 80% of X's staff and the company performed better. Private equity operating partners have built an entire arbitrage strategy around this: buy bloated companies, squeeze out the fat, pocket the spread. That's it.

Just as 80% of employees do nothing, 80% of tokens today do nothing.

People beget more people; tokens beget more tokens. Loop is the new land grab.

IV. 100X token is the new 10X engineer

Software's original promise was: write once, run cheaply forever, with no one watching. AI broke that promise — once software can do everything, nothing it does is stable or predictable anymore.

Tokens behave like a workforce. And once you start seeing tokens as employees, AI's promises begin to crumble:

  • "Tokens are more accurate than people" — but only if the prompt is right.
  • "Tokens are faster than people" — but after 100 retries, speed means nothing.
  • "Tokens don't play office politics" — but they will burn tokens to expand their territory.
  • "Tokens don't quit" — but they don't survive model updates or session restarts.
  • "Tokens are trustworthy" — but when they err, they're confidently wrong and perfectly formatted.

There's only one thing AI truly beats humans at: scalability. Scaling a human workforce demands enormous effort in recruiting, onboarding, and attrition; scaling tokens happens in an instant. That's precisely why mismanaged tokens are so costly — and why you must find that 100X token approach and scale it.

10X engineers built the last generation of companies; 100X token will build the next.

A tiny fraction of employees can make everyone around them 10X more productive; tokens work the same way — for any specific task, there's always some piece of context that can slash AI's workload by orders of magnitude. 100X leverage from tokens is real.

On average, humans are cheaper than tokens; but good tokens at scale are cheaper than humans.

And management is what turns the former into the latter.

V. Hoarding context is the latest way humans protect their jobs

AI is running headlong into a massive political problem inside companies, and it's only getting worse: employees don't want to teach AI their secret sauce.

They're starting to realize: these systems aren't just here to "help them" or "make them more productive."

Look at Meta: employees holding company stock, with every incentive to make AI work, are furious that the company is using their context as training data. And this is at a tech giant... This conflict is the drama coming to every industry.

For centuries, tacit knowledge — the stuff that can't be articulated — has been workers' armor. Medieval guilds guarded their craft secrets fiercely. AI is the first technology that demands workers surrender all of this, and essentially all at once.

No one will freely train their own replacement.

The people sitting on 100X tokens are precisely the least motivated to share them. Emotionally, structurally, politically — the organization is fundamentally hostile to the very technology most critical to its future.

VI. Eval is the new OKR

The best way to manage a token workforce is the same as managing humans: define what "good" looks like.

The only breakout AI use case untainted by office politics is coding. It grows the pie and makes every engineer stronger.

The mechanism behind this is eval (automated scoring systems built for specific tasks to measure AI output quality). Today 99% of AI revenue comes from coding because coding has built-in eval: code either runs or it doesn't — no gray area.

Broader cross-domain AI use cases will only come online once someone builds the corresponding evals. Specific evals matter far more than teaching employees to write prompts or handing them a chat harness. With evals, AI will eat the parts of the economy that code could never touch.

The real craft of management is translating fuzzy human processes into code, expressing the qualitative in quantitative terms.

A company's eval suite will become its most valuable asset.

OKR was the key to maximizing human team output; eval will be the key for token teams — and this team can scale infinitely. Eval is the path to 100X token.

Moreover, no two companies will have the same evals. Eval will become core competitive advantage. An organization running generic evals with generic agents has no advantage whatsoever.

VII. The next trillion-dollar opportunity is companies that sell transformation

For years, enterprises have been spending on AI: usage contracts with foundation model providers, off-the-shelf AI applications, internal builds. But all of this obscures a brutal economic truth:

No one has actually made AI work reliably yet.

Silicon Valley believes so deeply in this failure that its latest obsession has become: short today's business models. "Neofirms," or "AI-native services," are getting funded to target the $21 trillion services spend in the knowledge economy. The logic: incumbents are so mired in their own politics and processes that they'll never transform themselves.

Neofirms may indeed create competitive pressure that forces "trad firms" to adopt AI. But the greatest AI assets still lie inside incumbents: already-proven differentiated processes, paired with existing distribution channels, ready to scale.

In fact, the next wave of biggest businesses won't eat existing services spend — they'll sell a new kind of incremental service to existing players: "AI transformation companies" will be 10X the size of any neofirm.

"Transformation" sounds like a one-time project. But hidden inside is a Jevons paradox (the more efficient you get, the more you consume): every use case an organization deploys spawns ten new ones; the deeper a company goes into AI, the more transformation services it consumes, and the frontier of possibility pushes forward every day. Continuous AI transformation will become the only way to stay in the game.

Think about Palantir. On paper, it's the most disruptable company in software: a $500 billion market cap (note: Palantir has since fallen to over $300 billion), building custom applications for enterprises by hand. By the logic that has made SaaS nearly uninvestable, Palantir should have zeroed out before ServiceNow.

But Palantir hasn't, because it never sold software — it sold transformation.

Except transformation itself is no longer what Palantir used to do. In an AI-first world, it's no longer just ontology, custom software, and the occasional custom prompt. The real work is building evals, minimizing token spend, and understanding a business deeply enough to encode it.

Encoding what makes each company unique into agents will become the decade's greatest economic task.

VIII. It's time to manage

Every phase of the AI boom has had its guiding slogan.

First "sell picks and shovels in the gold rush" — so we built infrastructure. Then "sell Service-as-a-Software" — so we built neofirms. Now infrastructure is abundant, services are abundant, and what matters is making the trains run on time.

It's time to go inside enterprises: find those 100X tokens, document the loops that work, and redirect the intelligence being wasted at scale to where it belongs.

Humans just became cheaper than software. But whether human or software, someone still has to tell them what to do.

Finally, if this article resonated with you or you'd like to discuss anything with us, feel free to reach out: