Three Types of AI Applications
A financial task can cost $1.26 to run on a frontier model or $0.02 on a cheaper model. A 63x difference for the same task.
That is the example Rogo gives for why it built a model broker: they look at the work being requested, decide how much intelligence it actually requires, then send it to the appropriate model. A simple company profile does not need the same machinery as a complicated M&A analysis. Sending both through the most expensive model is wasteful.
We have spent much of the last few years asking which models will win, but I think that is one level too low. The more useful question for an AI application is: what should it own?
There are three layers in an AI application. The harness understands the job: context, memory, tools, permissions, workflows, evaluations and the sequence of actions required to get something done. The model supplies intelligence. The compute powers the model in converting context to output. Today there are three economically stable ways to assemble those layers.
| Architecture | Harness | Model | Compute | What makes it work |
|---|---|---|---|---|
| Intelligence factories | Own | Own | Control | Enormous scale across a very broad distribution of tasks |
| Specialists | Own | Own | Rent | Deep ownership of a coherent job or role where learning generalizes across customers |
| Work operating systems | Own | Rent | Rent | Broad ownership of an environment where many jobs depend on changing state and context |
The distinction between the last two is especially important. A specialist tries to own a job across many environments and often looks like a particular role at a company. A work operating system tries to own an environment across many jobs and often looks like much of the work done at a company. Both are applications. Both own a user interface, a harness and the relationship with the end user. The difference is where they draw the boundary around the work they want to own and where their proprietary learning compounds.
For a specialist, learning should increasingly become generalizable capability: something the system learns once and can apply across thousands of customers and millions of future tasks. For a work operating system, much of the valuable knowledge is stateful: what is true for this company in this workflow done for this customer at this moment. One belongs increasingly in the model and the other belongs increasingly in the harness.
At enormous scale, the application becomes an intelligence factory
The first architecture is the easiest to see because its products are already everywhere. Companies like OpenAI, Anthropic, and xAI are measuring their future committed compute footprints in the tens of GW, or enough power for tens of millions of homes. The physical scale is absurd because the economic scale is equally absurd. I put ChatGPT, Gemini, Grok and the largest general-purpose consumer-oriented AI products in this first category. These applications need to control the architecture, procurement, utilization and long-term supply of compute that underpins them since, at this scale, compute cannot remain a line item purchased above someone else’s margin.
A broad consumer-scale assistant has enormous task diversity that requires a generalist model and permits extraordinary infrastructure utilization. Every optimization in quantization, batching, networking, memory, chip design or model architecture can be amortized across billions of interactions.
The model, infrastructure and product become one machine out of necessity, but the danger is equally obvious. These are capital-intensive businesses whose infrastructure has to remain full. Owning the economics of a factory is wonderful when the factory is busy and not so great when the machines are idle.
A specialist owns one job across many environments
The second architecture starts with the observation that a general model knows far more than most applications need.
A customer-service agent does not need to write a screenplay, prove a topology theorem or explain nineteenth-century French politics. It needs to identify what a customer wants, understand company policy, execute reliably and escalate appropriately. Giving up capability outside that distribution can be valuable if the remaining model becomes cheaper, faster and better at the work that matters.
Decagon is probably the cleanest current example. It began by relying on foundation models. By March 2026, the company said more than 80% of its model traffic ran through models it had trained itself. Its system uses separate specialized models for functions including workflow execution, hallucination detection and understanding when speech has ended.
These transitions demonstrate that the choice around model ownership should be a consequence of the workload served.
The attractive specialist applications have a dense task distribution. Millions of interactions keep asking the system to perform variations of the same underlying job. The application observes failures and how humans correct them, making its evals more precise. The useful thing about those observations is that much of the learning transfers across customers. If an application discovers a better way to determine whether a customer has finished speaking, that insight can improve thousands of deployments. In other words, these applications rely heavily on tasks with generalizable knowledge.
Of course, a specialist still needs state, but the changing state is usually bounded by a relatively stable job. A legal application needs the matter, documents, relevant precedents and permissions. A support application needs the customer history, company policy and current status of an order. The company can build a very deep understanding of how the underlying job should be performed, then feed the relevant customer-specific context into that capability at runtime. The skill in the task itself dominates any given state.
I think the real moat in the second architecture is therefore broader than the model itself. It is the machinery for turning repeated work into reusable capability: proprietary task data, environments, evaluations and feedback loops, combined with a product that is purpose-built around the role. Harvey should understand how lawyers research, review, draft and collaborate. Specialization should run through both the model and the harness.
In essence, we are answering: Does learning from one customer’s work make the application meaningfully better at doing the same job for the next customer? If yes, then owning the model can improve product quality and gross margin simultaneously, while a specialized harness can make the entire application better suited to that role.
A work operating system owns one environment across many jobs
The third architecture starts from almost the opposite problem. These companies do not need one model to become exceptional at one bounded job. They need to perform many different jobs inside an environment that is constantly changing.
An investment firm is a useful example. One person might research a company, update a financial model, prepare a presentation, search internal documents, draft an email and coordinate a diligence process in the same afternoon. Another person in the same firm may spend the day doing a completely different set of tasks. The application wins by understanding the environment connecting those activities rather than becoming world-class at any single one of them.
The architecture looks almost exactly like an operating system allocating work across processors: understand the workload, supply the necessary context, pick enough capability and avoid paying for more intelligence than the task needs. Its deeper asset is the institutional layer underneath that routing: a representation of the firm’s data, precedents, workflows, people and history that persists even as the best underlying model changes.
The work operating system therefore needs to know something different than the specialist. Who is involved and what are their permissions? Which systems have the relevant information and are authoritative? What does complete mean? What does this team mean when they say “do it the usual way”? How does something that happened in finance affect what needs to happen in another department like sales, procurement or operations?
Much of that knowledge should never be baked into model weights since it changes too quickly and is contextually different by customer. It describes a current state rather than a reusable skill or capability. A model can learn how an AP process generally works, but the fact that invoice 4278 is awaiting approval from Sarah, that the supplier changed bank accounts yesterday and that the CFO suspended payments over $100,000 this morning is a different kind of knowledge that lives at inference time.
The model does not necessarily get smarter, but rather the system around it gets smarter about what the model should know and do. This is what I am calling stateful knowledge. If the application owns that representation, the underlying model(s) can become surprisingly interchangeable. When new models arrive, the application evaluates it on several hundred workflows, routes new appropriate tasks toward it if any, and keeps the customer’s memory, permissions, history and procedures exactly where they were before.
Generalizable knowledge belongs closer to the model; stateful knowledge belongs closer to the harness
This gives us a cleaner way to distinguish the second and third architectures. We can think of two kinds of complexity. The first is task entropy (how varied are the jobs the system needs to perform?) and the second is state entropy (how much does the info, workflow and environment shift from one situation to the next?).
A specialist tends to have lower task entropy. It may deal with enormous contextual variation, but the job itself remains recognizable. Customer support still resolves customer issues. Legal work still involves a related set of research, analysis, review and drafting workflows. Because the job remains stable, learning can compound across instances. The company gets repeated attempts at essentially the same game.
A work operating system has much higher task entropy and often much higher state entropy. The user may research a company at 9:00, update a spreadsheet at 10:00, prepare a presentation at noon and negotiate a contract in the afternoon. Each task touches different systems, requires different capabilities and depends on a changing set of facts. There is much less reason to force all of that through one proprietary specialist model.
The frontier labs and open-source ones behind them are already spending billions of dollars improving general reasoning, vision, coding, tool use and long-context performance. Renting that research program is actually a wonderful deal and guides our rough rule:
| Specialist | Work operating system | |
|---|---|---|
| What it tries to own | A job or role across many environments | An environment across many jobs and roles |
| Dominant proprietary knowledge | Generalizable skill | Stateful context |
| Task entropy | Lower | Higher |
| State entropy | Usually bounded by the job | High and constantly changing |
| Where learning compounds | Specialized model and role-specific harness | Data model, memory and orchestration |
| Model substitutability | Lower | Higher |
| Core question | Can we build a materially better application for this job? | Can we understand the environment well enough to perform many jobs? |
Of course companies can embody some of both. Take Harvey. Its legal reasoning layer is a specialist problem. Understanding statutes, interpreting clauses and reasoning through precedent can generalize across matters, so some of that learning belongs in a specialized model. Its institutional layer contains a firm’s documents, matter history, precedents, preferred language and permissions, which are stateful and belong in the harness.
But the important point is that Harvey is still a specialist application if it draws its product boundary around the lawyer and legal work. The presence of stateful knowledge does not make something a Type 3 application. What matters is the nature and scope of the work the application is trying to own.
The real question is when depth beats breadth
Both Type 2 and Type 3 are application architectures, which means both need to own a meaningful end-user relationship. A user is unlikely to maintain two AI applications that both claim to be the primary place where the same work gets done. If a work operating system can perform the specialist’s core job just as well while also understanding the surrounding environment, the specialist has a problem.
But Type 2 can still exist because there are jobs where the benefits of specialization are large enough to overcome the convenience and context advantages of breadth.
The first condition is that the job itself needs to be large enough to constitute a workspace. Software engineering is not one feature inside a knowledge worker’s day for an engineer; it is most of the day. The same can be true of legal work for a lawyer or customer service for an agent. These jobs contain many subtasks, but they sit inside a coherent function and can support a dedicated application because the user spends enough time there.
The second condition is that depth needs to create a material product advantage. The specialist has to be meaningfully more accurate, reliable, efficient or natural for the role. If the broad work operating system is 95% as good at a task the user performs occasionally, breadth probably wins. If the specialist is dramatically better at the work that consumes most of the user’s day, asking that user to work in a specialized application is not much of a burden.
Third, learning from the job has to generalize across customers. Harvey can spread the cost of learning legal reasoning across thousands of lawyers and matters. Decagon can turn millions of support interactions into a better support product. A work OS serving one company’s heterogeneous workflows may understand that company much more deeply, but it may never see enough repetitions of a particular function to move down the same learning curve.
Finally, the state required to perform the job needs to be sufficiently bounded or accessible. A legal application does not need to understand everything happening inside the corporation; it needs the state relevant to legal work. If that context can be pulled from the appropriate systems and documents, owning the rest of the enterprise provides less incremental advantage.
In simple terms, Type 2 earns its independence when economies of specialization exceed economies of scope and context. Type 3 wins under the opposite conditions.
The best companies own the bottleneck
In evaluating an AI application company, the question I come back to is what does production teach their system?
Does it teach a generalizable skill that makes the application better at a coherent job across every customer? If so, I want to understand whether that job is large enough to own the user, whether specialized model and product development can create a durable performance advantage, and whether the relevant state can remain sufficiently bounded.
Or does production teach the system more about a changing environment? If so, I want to know whether the company is actually building a persistent representation of the customer’s work across roles and systems: memory, permissions, relationships, workflows, state and history.
Those questions lead to three different conceptions of AI applications. The first type owns a broad consumer surface and integrates downward because its scale demands it. The second owns a role and integrates downward into the model because specialization compounds. The third owns an environment and keeps the model flexible because context compounds.
All three can create very large companies, but they get there through different economies.
The precise boundaries here will move and blur at times, but every AI application must rent the commodity portion of the task and own what bottlenecks their customer. When learning from repeated work generalizes, the model can become a bottleneck worth owning. When scale becomes extraordinary, compute can become one too.
The interesting question is not how much of the stack a company owns. It is where learning compounds in its workload, and whether the application owns the layer and the user relationship where that advantage accrues.