Generated by Codex with GPT 5.6 Sol XHigh
The Pragmatic Engineer surfaced Ramp’s internal coding system in its August 25 deep dive, “Why Ramp built its own in-house coding agent, Inspect.” The headline number is striking: Ramp says Inspect now originates 75% of the pull requests merged at the company. The more important result is architectural. Inspect is not a better foundation model. It is a model-agnostic agent wrapped in a remote developer environment, connected to Ramp’s internal systems, and required to verify its work with the same tools an engineer would use.
The environment is the product
Most coding agents begin with a model, a repository checkout, and a set of shell tools. Inspect begins with a reproducible cloud workstation. Each session gets an isolated sandbox containing Ramp’s development services, including databases, queues, workflow infrastructure, Chromium, and a browser-accessible VS Code instance. The Pragmatic Engineer reports that a fully provisioned environment starts in under five seconds.
That remote design removes two practical limits of local agents. Engineers can launch many sessions in parallel without exhausting a laptop, and Ramp can maintain one centrally configured environment instead of asking every user and agent to reproduce a complicated setup. A session can begin on the web or in Slack, continue through a Chrome extension, and hand control to a person in VS Code. Sessions are also multiplayer: colleagues can join the same running workspace rather than reconstructing its state from a pull request or chat transcript.
Ramp did not build the low-level coding harness or the models. Inspect uses the open-source OpenCode harness and can switch among frontier models. Ramp concentrated its effort on the layer that generic vendors cannot know in advance: how its repositories run, which internal services matter, and how the company decides that a change is ready.
Closing the verification loop
That internal context is useful because Inspect can do more than edit files. For backend work, it can run tests, inspect telemetry, query feature flags, and use a sanitized read-only production database to investigate mismatches between data and business logic. For frontend changes, it can start the application, drive a browser, and return screenshots and live previews. Ramp’s own account of the system lists integrations with GitHub, Slack, Buildkite, Sentry, Datadog, LaunchDarkly, and Braintrust.
The distinction is easy to miss. A general coding agent may produce a plausible patch after reading source code. Inspect is designed to gather evidence that the patch works in Ramp’s actual environment. Its advantage therefore comes from permissions, integrations, and feedback loops as much as from model intelligence. When a stronger model arrives, Ramp can place it inside the same environment without rebuilding those organizational connections.
The same platform supports work beyond prompted feature development. Engineers can ask Inspect from a Slack bug thread to investigate and raise a fix. An on-call assistant gathers incident context from production and observability systems. ReviewBuddy applies team-specific review rules. Testo explores the frontend and generates Playwright tests. Other agents analyze customer feedback, answer questions across internal data, or turn monitoring alerts into draft pull requests. The article says more than 200 internal agents now run on Inspect’s infrastructure.
Adoption is evidence, not a controlled experiment
Inspect’s reported adoption was fast. Its first version was a Chrome extension that let designers request small visual edits, but it saw limited use because engineers could make trivial changes faster themselves and non-engineers still needed local setup. The second version moved the entire environment to the cloud. Within two months, it originated about 60% of Ramp’s pull requests; the figure later rose to 75%, and the system passed one million sessions in July.
Those numbers show that people use the tool, but they do not by themselves measure productivity, correctness, or business value. Pull requests vary enormously in size, difficulty, and review cost. The figures are also company-reported, and the article does not provide a matched comparison with teams using commercial agents. Ramp’s private, production-derived coding benchmark is a more disciplined attempt to separate model behavior from harness effects, but even that benchmark is built from work already delegated to Inspect and represents Ramp’s own engineering distribution.
The organizational investment is substantial enough to qualify the success story. The core Inspect team is described as 5.5 people, while more than 150 Ramp engineers have contributed to its codebase. Every session is visible and collaborative internally, with no private-session opt-out. That openness spreads examples and makes agent work reviewable, but the wider system still carries real governance demands. An agent connected to source code, telemetry, databases, incident tools, and chat has a much larger blast radius than a local autocomplete tool. Read-only production access, sandbox isolation, scoped permissions, audit trails, and human review are part of the product, not optional hardening to add later.
A narrower case for building
The article challenges the usual “buy rather than build” rule for developer tools, but the lesson is narrower than every company needing its own coding agent. Ramp avoided rebuilding commodity layers. It reused an open harness, rented sandbox infrastructure, and swapped among external models while owning the company-specific environment and verification logic.
That boundary is the durable insight. General-purpose agents will keep improving, and vendors will absorb popular features. They still cannot arrive with safe access to a company’s internal data, production conventions, undocumented workflows, and definition of done. For organizations with enough scale, distinctive infrastructure, and the engineering capacity to secure and maintain it, that context layer can justify an internal platform. For everyone else, Inspect is less a blueprint to copy wholesale than a checklist for evaluating tools: can the agent reproduce the environment, obtain the right context, verify outcomes, expose its work to collaborators, and preserve a human-controlled review boundary?
Inspect suggests that the competitive unit in AI-assisted engineering is shifting. The model writes the patch, but the surrounding system determines whether the patch is grounded, testable, governable, and useful.