Generated by Codex with GPT-5
Techmeme surfaced Sarah Guo’s June 10, 2026 essay in its Techmeme item on The Untrainable. The original piece is The Untrainable, published on Sarah Guo’s Substack.
Guo’s essay is a useful answer to a question that hangs over almost every AI application company: if frontier models keep improving, what remains defensible above the model layer? The pessimistic version is simple. Models absorb tasks, wrappers collapse, and value concentrates in compute, chips, and frontier labs. Guo argues that this view is only half right. The measurable parts of work are being eaten quickly, but the most valuable work often depends on private data, private judgment, customer permission, liability, workflow change, and slow institutional trust.
That distinction makes the essay more interesting than another “AI apps versus model labs” argument. It does not claim that application companies are safe because models will stall. It assumes the opposite. Intelligence gets cheaper, benchmarks get saturated, thin wrappers get absorbed, and public tasks fall toward commodity pricing. The defensible zone is not where models are weak. It is where correctness cannot be established from the outside.
Benchmarks Mark The Commodity Frontier
The essay starts with software because software is the case that makes AI investors most anxious. Coding agents have improved at a pace that makes old arguments feel stale almost immediately. A system that looked limited in 2024 can look real by 2026, and the visible jump in benchmark performance makes it tempting to say that software engineering has simply been eaten.
Guo pushes back by separating code generation from engineering work. She cites research across more than 100,000 developers showing that modern coding agents increased the amount of code written far more than the amount of code that actually shipped. That gap is the point. Writing code is becoming cheap. Getting the right change through a real codebase, with real tests, real deployment constraints, old assumptions, hidden dependencies, ownership boundaries, and long-term maintenance costs, remains slower and more human.
Coding agents advanced early because software has unusually cheap verification. A compiler can reject invalid code. A test suite can catch regressions. A benchmark can score a patch. Once a task has a cheap verifier, the task becomes trainable. Models, scaffolds, and agents can grind against the check until performance rises.
But not all correctness is like that. A green test suite does not prove that a change belongs in a decade-old system. It does not prove that a new abstraction fits the product, that the deploy path is safe, or that the organization will be able to live with the decision. Some systems are trusted because they have survived years of actual use, not because they passed a fresh benchmark. The world itself is the verifier, and the world does not run faster just because the model does.
This is the essay’s first practical filter. If a task has public inputs and cheap public evaluation, the model ecosystem will eventually train against it. Benchmarks are useful, but they also advertise which work is becoming legible enough to commoditize.
Private Correctness
Guo’s core distinction is between work whose answer can be checked publicly and work whose answer can only be checked inside a private environment. A generic question answered by a generic model has little durable value. A model reasoning over a company’s data, with that company’s systems, constraints, risk tolerance, customers, and definition of success, can be much more valuable.
That leads to a rough map of AI work. Saturated work with public answers becomes commodity tokens. Frontier work with public answers is where labs have the advantage because they can train directly against visible evaluations. The interesting corner is frontier work with private correctness: tasks where the answer is only knowable inside a company, profession, or workflow.
The private part matters as much as the frontier part. A better model does not automatically get access to a bank’s production systems, a hospital’s clinical workflow, a law firm’s client files, or a customer’s support history. It does not sign the contract, pass the security review, accept liability, or build user habit. Intelligence helps, but permission and accountability decide whether the system can act.
That is why Guo calls the best territory “untrainable.” The label does not mean models can never improve at the task. It means the real feedback loop is not sitting in a public dataset waiting to be consumed. The truth is embedded in private systems and human judgment. A lab can improve the base model, but it cannot simply scrape the customer’s definition of a good answer.
Why Applications Still Matter
The essay’s defense of application companies is not sentimental. Guo is clear that many wrapper companies are weak. If a product adds a thin interface around a general model, does not own important private context, and can be replaced when the lab exposes the same feature, it is exposed. The labs are actively pulling retrieval, routing, tool use, reasoning policy, and other wrapper logic into the model layer.
The stronger application businesses do something less glamorous. They get trusted inside a workflow, organize private data so a model can use it, wire the model into tools that can act, and help the customer change how work actually happens. That integration work is slow, domain-specific, and hard to copy from the outside.
Law is one of Guo’s examples. A large law firm’s M&A practice is not one generic document task. It is a chain of matter-specific work: nondisclosure agreements, term sheets, diligence, purchase agreements, ancillary documents, closing checklists, client expectations, partner judgment, and firm-level coordination across many active deals. The important signal lives in how the matter flows, not in one associate’s isolated prompt.
Healthcare has a similar shape. A foundation lab may produce a strong medical model, but that does not mean it owns the physician’s habit, the hospital’s workflow, or the professional authority to define safe answers in context. A product such as OpenEvidence matters because users have already allowed it into the daily loop. The model is only one part of that adoption.
Customer support, legal automation, software agents, and inference serving all show related patterns. Sierra can price around resolved support outcomes because it works inside the customer’s definition of resolution. Devin can offer performance guarantees only where it has enough access to observe the work. AI-native companies may care less about raw token price than about reliability, traffic behavior, and scarce compute access under real load. In each case, the value is not merely model intelligence. It is trust under operating conditions.
Owning The Evaluation
The most important strategic idea in the essay is that defensible companies can earn the right to define what “good” means in a field. Public benchmarks matter while the task is outside the workflow. Once the work happens inside a customer’s system, the decisive evaluation becomes private: this company, this task, this risk tolerance, this user, this acceptable outcome.
That creates a path for application companies to build moats. They do not just use models. They collect the judgment that turns ambiguous work into repeatable standards. They learn what the customer accepts, what users override, what lawyers approve, what doctors trust, what support teams escalate, and what operations teams can actually maintain. Over time, those decisions become a domain-specific evaluation system.
This kind of evaluation authority is hard for a lab to create from the outside. It comes from adoption, not just capability. The senior lawyer, physician, engineer, operator, or customer-support leader has standing because the field already treats that person’s judgment as relevant. The product that captures and operationalizes that judgment can become more than an interface. It becomes part of how the field decides what counts as good work.
This also explains why outcome pricing matters. If a vendor can charge only when a customer issue is resolved or when a coding task meets an agreed performance threshold, the price itself becomes an evaluation. That is only possible when the vendor is embedded deeply enough to observe the outcome and trusted enough to take responsibility for it.
The Moat Keeps Moving
Guo does not present the untrainable zone as a place where companies can relax. The absorption frontier keeps rising. Every time the industry learns how to measure a task, that task becomes easier to train against and easier to commoditize. What was private and hard to score yesterday may become a benchmark tomorrow.
That makes the strategy dynamic. A company has to keep moving toward work that is still hard to evaluate from the outside. It has to keep deepening access, improving private feedback loops, building workflow trust, and expanding the scope of what it can safely automate. A narrow private task can support specialized models and strong economics. A general public task becomes a compute contest.
The warning for founders and investors is direct. If the product’s survival depends on out-training frontier labs on broad general tasks, the likely winner is whoever owns the most compute. If the product can get inside a valuable private workflow, define the outcome, and keep learning from customer-specific reality, it has a much stronger case.
The essay also points to a scarce input that models do not provide: intent. Models can execute, draft, search, reason, and generate options. They do not decide which new market should exist, which workflow is worth rebuilding, or which ugly enterprise problem deserves years of work. The ability to choose what to build remains outside the benchmark loop.
Takeaway
Techmeme was right to surface this essay because it gives a cleaner test for AI application value than the usual argument about whether models are “good enough.” The question is not whether frontier models will keep improving. They will. The question is whether a company owns something the model cannot learn from public evaluation.
A defensible AI business needs more than a prompt, a workflow wrapper, or a temporary model gap. It needs access to private context, permission to act, a way to observe outcomes, user trust, domain judgment, and a role in defining what good work means. Those ingredients are slow and messy, but that is exactly why they matter.
The thin-wrapper layer is in danger. The untrainable layer is harder to attack because its value lives in history: customer relationships, embedded workflows, private data, accumulated judgments, and institutional accountability. The best AI companies above the model layer will be the ones that turn those messy realities into systems a model can use.