A Quiet Beginning, a Major Tectonic Shift
On September 10th, OpenAI released a single API. There was no flashy announcement, no new model name. It's a public beta called "Agents API," a rather unassuming name. But upon reading its contents, I felt this might be the most impactful announcement on how engineers work in the last few months.
In short, OpenAI has decided to lend out the entire "backend mechanism that powers Codex." The cumbersome foundational elements that have supported ChatGPT's "Work" function and the coding agent Codex—session management, context processing, and sandbox execution—are suddenly accessible through a single API call.
What is a "Harness"?
For those unfamiliar with the term, a "harness" is like a scaffolding for running agents. No matter how intelligent a model is, to use it for actual tasks, it needs a mechanism to maintain conversational context, a mechanism to invoke tools, a mechanism to assign tasks to multiple sub-agents, and a sandbox for safely executing code.
Until now, many teams built this "scaffolding" themselves. They wrote prompt chains, managed tool call logic, and implemented their own state-saving mechanisms to prevent session interruptions. While seemingly mundane, this was the most labor-intensive part when creating agents that would run for extended periods in production.
Made up of four components
According to OpenAI's official documentation, the Agents API is composed of four concepts: an "agent" that combines the model, tools, and MCP server; an "environment," or sandbox, for file access and command execution; a "session" that maintains state between turns; and "events and items" representing input and output themselves.
The sandbox selection is also flexibly designed. You can use a sandbox hosted by OpenAI, or you can set up your own `codex exec-server` on your infrastructure and connect. It also integrates with sandboxes from partner companies such as Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel from the start. The design philosophy is that developers only choose "where to run," and OpenAI handles the rest.
The Invention of Context "Compression"
Personally, what I found most interesting was the automatic context compression mechanism for long-running sessions. When an agent works for many hours, the conversation history and work logs grow exponentially, exceeding the model's processing capacity. The Agents API automatically summarizes and compresses the context up to that point as the session approaches its limit, leaving only the information necessary for the agent to continue working.
Another mechanism is "tool search." By loading only the definitions of tools likely to be used each time, it maintains the model cache while minimizing token consumption. Furthermore, using "programmatic tool calls," tools can be executed in parallel, and large amounts of data can be filtered code-wise before returning only the necessary results to the context. While this may seem like a minor feature, it directly addresses the biggest cause of the high costs associated with deploying agents in production.
Sandboxes are selectable, but data cannot leave the US
On the other hand, there are some concerning limitations. Currently, the Agents API's data residency is limited to the US and does not support Zero Data Retention (a mode that does not retain any data). Even if a company sets up its own sandbox in any country, the control plane itself remains within the US. This is likely to be a barrier to adoption for regulated industries and companies outside the US.
Voices from Companies That Have Actually Used It
The article also includes testimonials from companies that have already implemented the API. The head of engineering at insurance tech company WithCoverage commented that while they previously wrote their own prompt chain and tool call management, the Agents API allowed them to rethink the design of their complex, multi-stage workflows. One company reported a 60% reduction in costs and improved latency after migrating their case review workflow. While not flashy figures, these are reliable reports of cost savings in real-world operation.
From "Model Intelligence" to "Infrastructure Maturity" as the Competition Shifts
This API suggests that the competitive landscape for frontier AI is gradually shifting. This marks a shift from competing on model benchmark scores to competing on "how mature the foundation for running agents stably for extended periods has been." The pricing structure is also simple: there are no additional charges for the Agents API itself; you are only billed for the tokens and tools you use.
Things Engineers Should Consider
For teams building their own agent infrastructure, this announcement provides material for seriously considering whether to switch or continue developing in-house. The ability to externalize parts that subtly consume time with each feature addition, such as session management and context compression, is particularly significant. However, the US-only data residency restriction is a significant obstacle for teams running agents in overseas locations, including Japan. For the time being, it seems realistic to monitor how this restriction is eased and try implementing it via the partner sandbox.