What actually changed from ChatGPT wrappers to vertical AI, and where I think the current thesis still needs more scrutiny
A post from Garry Tan caught my attention recently.
Someone who had attended YC Demo Day observed that, outside of hardware and physical-world companies, a lot of startups seemed to be building some form of domain-specific harness.
Garry replied:
“Either you die a system of record or you live long enough to become a domain-specific harness.”
It is a good line. It also describes a real change in how AI applications are being built.
But it left me with a question.
Not that long ago, investors were warning founders against building “ChatGPT wrappers.” The concern was that if most of the intelligence came from OpenAI or Anthropic, and your company added a prompt, some RAG and a user interface, there wasn’t much left to defend.
Now we seem much more comfortable funding companies that sit above foundation models and apply them to a specific domain or workflow.
The technology is clearly different.
I am less convinced that the economics are always as different as the terminology suggests.
That is worth unpacking, particularly if you are investing in vertical AI.
A harness really is more than a wrapper
I don’t think it is useful to call every application built on a foundation model a wrapper.
The early version was often fairly simple:
User → prompt → model → response.
Maybe there was retrieval. Maybe some company documents were added. The better products had useful interfaces and integrations. But much of the underlying intelligence still belonged to the model provider.
What people increasingly call a harness is more substantial.
The application may maintain state, collect context from several systems, decide what tools to use, call APIs, execute actions, verify the result, recover from failures, respect permissions, escalate decisions to a person, maintain an audit trail and measure whether the workflow actually succeeded.
So yes, this is materially different from putting a nice interface around ChatGPT.
The problem is that technical complexity and defensibility are not the same thing.
You can spend a lot of time building orchestration, memory, tools, evaluations, permissions and exception handling and still end up with a product that somebody else can reproduce.
That is where I think the current vertical-AI discussion becomes more interesting.
The incumbent has more advantages than we sometimes acknowledge
Take Salesforce.
A startup building an AI layer for sales, service or CRM may have a better user experience and may move much faster than Salesforce.
But Salesforce starts with some things that are difficult for a startup to acquire: the customer relationship, historical data, permissions, workflow logic, integrations, administrative infrastructure and enterprise trust.
And Salesforce isn’t waiting for startups to take the agent layer.
A few days later it announced additional integrations that connect its data, workflows and business logic with Google’s Gemini infrastructure and agents.
The exact products will change. The strategic direction is more important.
The system of record is moving into the execution layer.
Healthcare makes the point even more clearly for me.
Epic says more than 85% of its customers now use Epic AI.
For an AI startup selling into health systems, this isn’t an abstract future threat.
The incumbent already has the EHR integration. It has the data. It has the user accounts. It understands the permission structure. It has years of procurement history with the customer.
Now it can add intelligence and agency.
That makes one question increasingly important when I look at vertical AI companies:
If the incumbent ships 80% of your functionality next year, what is left?
I don’t think “our prompts are better” is a good answer.
Neither is “we use multiple models.”
Even integrations are a weaker answer than they were a few years ago. Integration work certainly creates friction, but standards are improving and models themselves are becoming better at interacting with software.
The founder needs a more durable answer.
Where I think the startup opportunity is real
None of this means that incumbents win by default.
In fact, market data suggests otherwise.
Menlo also estimates that startups captured 63% of AI application revenue in its dataset.
So clearly startups can take share from incumbents.
The question is where.
I think one answer is that a system of record often owns a record without owning the entire job.
Salesforce contains the opportunity. It does not necessarily own everything required to close the sale.
Epic contains the patient record. It does not necessarily own every interaction involved in getting care authorized, scheduled, delivered and reimbursed.
DocuSign contains the signed agreement. It doesn’t own the complete legal process that produced it.
Actual enterprise work is messy.
It crosses applications, organizations and sometimes companies. It involves email, PDFs, spreadsheets, phone calls, legacy databases and, in healthcare, still an uncomfortable amount of fax.
That fragmentation creates room for a startup.
Andreessen Horowitz recently made a similar argument in an essay called The Incumbents Are Coming. The piece makes the bullish case for systems of record, but argues that vertical AI companies can still win by becoming better at performing the actual job, particularly where the work crosses systems and expert judgment matters.
That framing makes much more sense to me than simply saying vertical AI wins because it is specialized.
Specialization by itself isn’t enough.
A startup needs to own something meaningful that the incumbent doesn’t naturally own.
The bigger opportunity may be selling the work rather than the software
This is where vertical AI becomes much more interesting than another SaaS cycle.
Traditional SaaS usually gives a person better software for doing a job.
The more ambitious AI-native company can eventually say:
Give us the job.
Consider a workflow that costs a company several million dollars a year in internal staff or outsourced services.
A SaaS vendor might sell software that improves the team’s productivity by 20%.
An AI-native company might eventually say:
“We will execute 70% of this process. Your team handles the exceptions.”
That is a different economic proposition.
The startup is no longer competing only with another software license. It is competing with labor, BPO spend, consulting spend and internal operating cost.
It also starts taking responsibility for the outcome.
I think that distinction matters much more than whether the product is called a copilot, agent or harness.
A company that recommends what an employee should do has one level of value.
A company trusted to actually do the work has another.
Domain expertise becomes more important, not less
There is another part of this market I remain skeptical about.
It has never been easier to build an impressive vertical-AI demo.
A capable engineer can take a frontier model, add several tools and produce something in a few weeks that would have looked impossible three years ago.
That can create the impression that understanding the domain has become less important.
I think the opposite may be happening.
The model already knows an extraordinary amount about medicine, finance, insurance, law, sales and many other industries.
What it often doesn’t know is how a specific organization actually works.
It doesn’t automatically know which data people trust and which fields they ignore.
It doesn’t know that a workflow described in the SOP is not the workflow employees actually follow.
It doesn’t know which exception happens once a week and costs $50,000 when mishandled.
It doesn’t know when a clinician, underwriter or compliance officer is comfortable allowing an agent to act without approval.
It doesn’t know which payer continually rejects a particular document, which approval gets stuck with which team, or why an employee keeps exporting something to Excel before making a decision.
That knowledge is often the difference between a good demo and a production system.
And much of it doesn’t exist neatly in a database.
It lives in people.
That is why I am wary when founders enter complex industries with little operating experience and assume the model will supply the domain expertise.
The model can provide knowledge.
Operating judgment is something else.
The real proprietary data is probably not the documents
Almost every vertical-AI pitch eventually includes some version of “our proprietary data becomes a moat.”
Sometimes that is true.
Often it needs more examination.
Customer documents aren’t necessarily the startup’s data. Contracts may restrict how information can be used. Data from one customer may not transfer cleanly to another. And general models are getting better at acquiring domain knowledge anyway.
The more interesting dataset, in my view, is the one created while the work gets done:
Situation → decision → action → correction → outcome.
What did the agent recommend?
What did the expert change?
Why did they change it?
What action was eventually taken?
Did it work?
That feedback loop is potentially much harder to reproduce than a collection of PDFs.
Suppose two companies use the same underlying foundation model.
One has processed 50,000 real workflows and knows where the model fails, when a human needs to intervene, what actions lead to good outcomes and how performance varies across edge cases.
The other has better prompts.
I know which company I would rather underwrite.
But even here, I would ask a very specific question:
What is materially better for customer 100 because customers 1 through 99 existed?
If the founder cannot explain that precisely, “data moat” may just be another phrase.
Evals may be a bigger moat than prompts
This also changes how I think about evaluation.
For a consequential vertical workflow, the important question isn’t whether the model performs well on a public benchmark.
It is whether the company understands when its system is right, when it is wrong and when it shouldn’t act at all.
Did the claim get paid?
Did the prior authorization get approved?
Did the contract contain the correct terms?
Did the customer issue get resolved?
Did the patient actually receive the care?
Those are operational outcomes.
Anthropic’s guidance on evaluating production agents emphasizes feedback from the environment, checkpoints and ground truth rather than simply making the agent more complicated.
Over time, the company that has seen thousands of real failures may build something valuable that isn’t visible in the demo: a very good understanding of what not to automate.
That matters a lot in healthcare, finance and other industries where one bad action can be far more expensive than 100 successful ones are valuable.
Open standards cut both ways
Healthcare is a useful example of another structural change.
That can make it easier for startups to work across systems that were historically difficult to access.
But it also reduces the integration advantage over time.
Open standards are good for challengers because they make incumbent data more reachable.
They are also good for incumbents and general-purpose agents because everyone can increasingly interact with the same systems.
Again, the integration itself isn’t necessarily the moat.
What you do with it matters more.
What I am actually looking for
When I look at a vertical-AI company now, these are the questions I find myself asking.
What happens if the incumbent gives away most of this functionality?
Does the startup own a feature, a workflow or an outcome?
Does the job naturally cross multiple systems, or does one incumbent already control most of it?
Where does the domain expertise come from?
What does the company know after 10,000 workflows that it did not know after 100?
Does better model capability strengthen this business or make the product easier to reproduce?
Does the company capture expert corrections and real outcomes?
And perhaps most importantly:
Is this becoming somewhere the work actually happens?
Sequoia gave founders an interesting piece of advice at AI Ascent this year: “Build moats from the customer back.” It also described models, tools and harnesses as the three components coming together in the current agent wave.
I think the first observation matters more than the second.
The technology stack tells us what can now be built.
It doesn’t tell us where durable value will accrue.
The harness may eventually become the system of record
There is one scenario where I think this becomes especially interesting.
A startup begins by sitting above existing systems.
It executes the work.
Because it executes the work, it sees decisions and exceptions.
Because it sees those decisions, it develops better judgment about the workflow.
Eventually users start asking the new system what happened, why it happened and what should happen next.
At some point, the most important operational history may no longer be the static record inside the old application.
It may be the sequence of actions and decisions created by the agent.
That is when the startup starts moving from an interface or feature to a system of action.
And a system of action that creates the authoritative history of work may eventually become a new kind of system of record.
I think that is a much more compelling end state than simply putting AI on top of existing software.
Where I land
I think domain-specific harnesses are a real and important architectural shift.
I don’t think “we built a domain-specific harness” is a defensibility argument.
Those are different statements.
A few years ago, the weak pitch was:
“We built a vertical interface around GPT.”
Today the weak pitch may become:
“We built a sophisticated agentic workflow around GPT.”
There is considerably more engineering in the second company. It may also deliver much more value.
But engineering complexity alone doesn’t make a company difficult to compete with.
The questions are still familiar.
Who owns the customer?
Who owns the workflow?
Who has the trust?
Who owns the outcome?
Who learns faster?
And who becomes harder to remove every year?
The vertical AI companies I am most interested in will certainly use harnesses.
They just won’t win because they have one.
They will win because that harness allows them to understand a job better, execute more of it, learn from every outcome and gradually become part of the operating infrastructure of the customer.
That is a much higher bar than a ChatGPT wrapper.
It is also a much higher bar than simply calling the wrapper a harness.
Sources and further reading
This essay was prompted by Garry Tan’s September 2026 post on systems of record and domain-specific harnesses, and draws on recent work and data from Anthropic, Salesforce, Epic, Menlo Ventures, Andreessen Horowitz, Sequoia Capital and CMS.
- Garry Tan: “Either you die a system of record or you live long enough to become a domain-specific harness.”
- Anthropic: Building Effective Agents
- Anthropic: Demystifying Evals for AI Agents
- Salesforce: Salesforce Introduces the Trusted Enterprise AI Harness
- Epic: Advocate Health, ECU Health See Early Wins with Agent Factory
- Epic: Real Results, Right Now: How Epic AI Is Reducing Costs, Improving Care, and Helping Patients
- Menlo Ventures: 2025: The State of Generative AI in the Enterprise
- Andreessen Horowitz: The Incumbents Are Coming
- Sequoia Capital: AI Ascent 2026
- Centers for Medicare & Medicaid Services: Interoperability and Prior Authorization Final Rule