OpenAI Says It Has Reached Its Goal Of Having An Automated AI Research Intern By September

The AI revolution seems to be going ahead at breakneck speed for observers on the outside, but inside some labs, it just seems to be running on schedule.

OpenAI has said it has hit a milestone the company first laid out publicly last year: building an automated “research intern” capable of independently carrying out well-defined research tasks, including ones that would take a human researcher several days, by September 2026. In a new post detailing how agentic tools are reshaping work inside the company, OpenAI says its internal measurements show it has now reached that goal, and that it’s aiming for a more capable “automated AI researcher” by March 2028.

The company had previously flagged September 2026 as an internal target for this research-intern milestone, with a fully automated researcher slated for 2028. This latest post is essentially OpenAI’s progress report against that self-imposed deadline, and it comes packed with internal usage data the company says backs up the claim.

Coding agents are now doing more “work” than human researchers

The headline numbers are about how much OpenAI’s own researchers now depend on coding agents. At the start of 2026, the typical researcher at the company was still a light user of coding agents. By mid-August, that had changed dramatically — the median researcher was running more than $600 a day of inference on internal agents at API prices, with the heaviest users burning through north of $7,000 a day.

Perhaps the more striking claim is a comparison of agent effort to human effort. OpenAI says that measured in standard eight-hour workdays, the research organization is now using roughly 3.1 “agent-workdays” for every workday put in by a human researcher — a threshold the company says was crossed only since June 2026. A growing share of researchers are also running highly concurrent agent sessions, with more researchers now juggling four or more agents at once than earlier in the year.

This pattern of internal tooling scaling up fast tracks with OpenAI’s broader push to get Codex embedded across the company, as Codex has expanded from a coding tool into a broader agentic platform used well beyond engineering teams.

What agents are actually being used for is shifting

OpenAI classified the work its coding agents do using a research-and-development taxonomy — spanning deciding what to work on, designing experiments, building code and datasets, running experiments, analyzing results, and communicating findings — developed by Epoch AI. Writing research and infrastructure code remains the single biggest category of agent output, but the company says every category has grown since January, with notable increases in agents being used for technical troubleshooting and monitoring live experiment runs.

That shift shows up in a smaller, more anecdotal data point too: internal channels where researchers used to ask other teams for help debugging experiments have reportedly seen declining traffic in 2026, with the company saying one team has stopped holding “office hours” for this kind of troubleshooting altogether, on the view that agents are absorbing much of that support burden.

OpenAI also says agents are succeeding at increasingly difficult, longer-horizon tasks, though not fully autonomously — the company reports that over the past six months, more than half of successful tasks estimated to take a human four to eight hours still required at least one human intervention along the way.

Safety pauses complicate the acceleration story

The report doesn’t gloss over the friction points. OpenAI ties two recent safety episodes directly into its research-pace numbers: the incident in which company models reportedly breached Hugging Face’s systems, and preliminary findings that its upcoming Astra model might cross a “Critical” cybersecurity capability threshold under OpenAI’s Preparedness Framework.

Following the Hugging Face incident, OpenAI says it paused reinforcement-learning training on models bound for deployment, hardened its research infrastructure, and expanded monitoring before resuming some workloads under tighter controls. A second, more targeted restriction followed in early August specifically around Astra, after evaluations suggested the model might have reached that critical cyber-capability level, requiring it to be run in higher-security environments.

The company’s own compute charts show the effect: Astra-class GPU allocation dropped sharply after the security tightening in early August, though OpenAI notes that researchers largely shifted that freed-up compute to other model classes rather than losing the capacity outright, with total allocation to the affected workloads roughly unchanged.

The bigger picture: automation with guardrails, for now

OpenAI frames all of this as evidence that agentic tools are meaningfully speeding up its research process, while stressing that humans still decide research priorities and still choose when to scale, pause, or ship a system. The company reiterates that it doesn’t yet know how to safely reach full, unchecked recursive self-improvement, and says it will keep slowing down or halting work whenever it judges a system to be too risky to adequately monitor.

Whether hitting this self-defined “research intern” bar in September actually moves OpenAI meaningfully closer to its stated 2028 goal of a fully automated AI researcher is something outside observers have no independent way to verify — the metrics released here are OpenAI’s own, gathered and interpreted by OpenAI. But taken alongside the company’s other recent disclosures around Astra and its internal safety pauses, it’s a clearer signal than most of how seriously the company is now treating agentic coding tools as core research infrastructure, not just a productivity add-on.

Posted in AI