The Task Economy Explained: Why AI Needs Vetted Experts

For the last couple of years, the AI industry has measured its own growth in tokens.

Public companies report monthly tokens processed as a headline metric. Analysts rank models by relative token volume. Management teams point to their token usage over time as proof of how seriously they’re investing in AI. It’s a clean number, and it does its job well, but it only tells you how much AI is being used, not what’s being done to make it better.

Benchmark general partner Everett Randle made that second, less visible market the center of a new piece titled “The Task Economy: Data Will Be the Next $1 Trillion Category.” His argument is that tokens measure usage and compute, not improvement. Improvement is a separate market: the market for the data that makes AI capable of expert-level work in the first place. He calls it the Task Economy, and he’s betting it becomes as central to the AI conversation as tokens are now.

We think he’s right. And we’d push the argument one step further: once this spend spreads past the frontier labs, the binding constraint won’t be data infrastructure. It’ll be finding enough qualified experts to actually do the work.

What a “Task” Actually Means Here

Randle’s use of “task” comes from reinforcement learning. A model is given a starting point and an environment to work in, it acts, and a reward signal or verifier scores how well it did. He’s careful to separate this from the older idea of data labeling — the bounding boxes and thumbs up/down ratings that defined the category a few years ago. What’s happening now is more demanding: real domain experts completing real professional work, graded against a rubric, so a model learns not just what to do but how to do it well.

His example is legal work. A model can absorb public case law from the open internet, but it can’t learn to practice law that way. You have to hand it a real task (review this contract, draft this argument), place it in a realistic environment, and grade the output the way a senior attorney would. Mercor, the company Randle’s fund backs, publishes exactly this kind of rubric for its corporate lawyer agent.

That definition is doing more work than it looks like. It quietly redefines “training data” as a specialized labor category: sourced, screened, and graded like a hiring pipeline rather than scraped like a dataset.

The Numbers Behind the Thesis

Randle backs the argument with several striking figures. OpenAI and Anthropic are reportedly scaling data spend 10x year over year. Leading AI application companies and enterprises in his network are moving toward individual task-related budgets north of $100 million. And Mercor, the platform he holds up as the clearest proof point, went from a $1 billion to a $2 billion annualized revenue run-rate in roughly four months, with expert hours worked on the platform growing exponentially alongside it.

What stands out to us isn’t the size of those numbers, but how concentrated the supply side still is. Nearly all of this demand is flowing through a small number of platforms. Some of that concentration is the usual stuff such as capital, first-mover advantage, direct relationships with the labs. But at least part of it, we’d argue, is simpler and more stubborn: sourcing and vetting genuine expertise at this speed is hard, and hard to scale.

Why This Has Stayed Under the Radar

Randle offers two reasons the Task Economy hasn’t drawn the attention inference has: spend has been concentrated inside frontier labs that don’t advertise how they improve their models, and there’s no simple public proxy for task volume the way there is for tokens.

Both are fair. We’d add a third. From the outside, this market looks like a data business: a more fundable, more legible story to tell investors. Underneath, it runs on something far less glamorous: recruiting, screening, and managing real experts at speed.

That’s a talent-operations problem inside a data-infrastructure market. Staffing problems don’t usually get the hype cycle that data does, even when they’re the actual constraint.

That constraint is also why Randle expects spend to broaden this year, from labs to AI application companies and enterprise teams outside the frontier. As more of them realize that a differentiated data strategy, not just access to an off-the-shelf model, is what separates their AI products from everyone else’s, they’ll need to build this capability themselves.

What This Means Once You’re Not Mercor

Here’s the part Randle’s piece doesn’t spend much time on, because it’s written for investors tracking a spending category, not for the enterprises about to enter it: most companies stepping into the Task Economy this year don’t have Mercor’s infrastructure. They don’t have an in-house pipeline for identifying, evaluating, and onboarding people with the right domain expertise, let alone doing it at the pace this category is moving.

Some of that need is genuinely specialized. Licensed professionals grading tasks in law or medicine are their own vertical, with their own credentialing requirements, and that’s not a gap you close with general staffing. But a lot of what’s actually being asked for doesn’t require a bar card or a medical license. It requires reviewers who understand a domain well enough to judge whether an agent’s output holds up, catch the edge cases, and flag what still needs a human sign-off.

Put in Randle’s terms, these reviewers are the verifier in the reinforcement loop, the human reward signal that tells the model whether it got the work right. That makes it a hiring and vetting problem long before it’s ever a machine learning problem. And it’s the layer where most enterprises are least prepared.

Where Persona Talent Fits In

This is a problem Persona Talent already solves. Persona’s model has never been about generic full-time hiring. It’s about matching specific, vetted people to specific work fast, with real screening behind every match, sourced from more than 1.4 million applicants a year across 98 countries.

That’s the layer of the Task Economy we think is most underserved right now: human-in-the-loop reviewers, evaluators, and ops-adjacent specialists who can support AI and automation teams by checking agent output, flagging errors, and handling the judgment calls that shouldn’t be automated yet. It’s an extension of the executive support, operations, marketing, and project management talent Persona has placed for years.

Judging whether an agent’s output is actually correct is a domain-competence screen, not a generic one, so the match matters more here than in ordinary staffing. That’s the part Persona is built for: a large top-of-funnel narrowed by rigorous, role-specific vetting, with people matched to the domain and seniority a given task demands. For the most specialized work, that pairs with a client-defined rubric. Persona supplies the vetted human judgment; the client defines what “correct” looks like.

As Randle puts it, “when we talk about AI, tasks will be king.” The companies that keep up will be the ones who can staff that work with people they actually trust.

Building out a review, evaluation, or human-in-the-loop function and need people you can trust fast? Talk to Persona Talent. We source vetted talent globally and can typically place someone qualified within 7 days.

Talk to our team and learn what makes Persona’s hiring process faster and more effective than traditional hiring methods.

"*" indicates required fields

Related Articles

BELAY vs. Boldly vs. Persona: Virtual Assistant Companies Compared
Show All Business articles

Sign up to receive regular insights on talent