Key Moments
Inside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj
Key Moments
OpenAI's new computer use agents are now faster than humans and cost 1/7th of previous models, enabling complex tasks like meal planning in minutes instead of hours, but questions remain about their ability to outperform expert human users.
Key Insights
Computer use agents can now perform tasks 7x faster than before, with new models like GPT-4.1 Soul costing 1/5th of Astra and 1/7th for computer use specifically.
The new Agents API allows developers to build on the same computer use capabilities previously only available in Codex and ChatGPT.
Dots, a new personal assistant product, offers each user a dedicated virtual Linux computer in the cloud, capable of running full desktop applications and web browsers.
The Decisions API, while not dependent on reasoning, performs inference in parallel and uses a smaller, faster model for tasks where complex, long-term reasoning isn't critical.
OpenAI's computer use now writes JavaScript code to execute multiple actions simultaneously, significantly improving speed and capability, and utilizes multimodal inputs like screenshots and accessibility features.
The new Responses API has been rewritten for maximum speed in Time To First Byte (TTFT) and Time To Last Byte (TTLB), with improved caching and pre-warming features for continuous agent operation.
Computer use agents achieve superhuman speed and cost efficiency
Computer use agents have advanced to the point where they are now faster than the average human for most tasks. The next frontier is achieving 'superhuman' performance, matching or exceeding expert human computer users. This advancement promises more interactive and immediate product experiences. This was not a certainty just a few months ago, but OpenAI has rapidly iterated, with internal teams embracing and developing the idea. The new GPT-4.1 Soul model is particularly notable for computer use, costing only one-fifth of Astra and a seventh for specific computer use tasks. The new Decisions API, while distinct, performs inference in parallel using a smaller model, making it fast but less suited for complex, long-term tasks, highlighting ongoing research into integrating different approaches.
New products and APIs empower developers and users
OpenAI DevDay saw several key announcements. Dots, a new personal assistant, provides each user with their own private virtual Linux computer in the cloud, enabling the execution of full desktop applications and web browsers. This is a significant shift from previous cloud browser or personal computer access. The Agents API now integrates computer use capabilities, allowing developers to build upon the same functionality present in Codex and ChatGPT. Demonstrations included App Shots, where context from user actions can be quickly input into AI models, and enhanced existing computer use features on Mac.
The power of dedicated cloud computing for AI agents
A core innovation is that each 'Dot' now has access to a dedicated virtual Linux computer in the cloud. This is a departure from traditional cloud access methods, offering a full computing environment. This capability is crucial because most software is designed for humans, and now AI agents can leverage these same programs. Users can delegate any task they perform on a computer to a Dot, from booking flights and shopping to more complex, time-consuming activities. An example cited was a user spending two hours customizing a meal delivery order, a task an agent completed in just 15 minutes, saving significant time and effort.
Evolution of computer use through advanced techniques
Computer use has evolved significantly, moving beyond basic task initiation to sophisticated error correction and iterative attempts. A key development is that computer use agents now often write code, such as JavaScript, to perform multiple actions simultaneously, drastically improving speed and capability. The interaction layer has also become more sophisticated, utilizing multimodal inputs like screenshots, accessibility features, and tools like Playwright. This allows models to 'see' entire pages or applications and execute complex sequences of actions, rather than just single steps.
Enhanced computer use through multimodal inputs and accessibility
The integration of multimodal inputs has been a major breakthrough. Previously, agents would struggle with tasks requiring scrolling, taking multiple screenshots, and re-evaluating. With improved accessibility features and direct DOM access, models can now perceive entire pages or applications at once. They can write code to perform multi-step actions efficiently. This also means that technologies designed for users with disabilities, like screen readers and accessibility representations, are proving highly beneficial for AI agents, providing rich context for tasks like identifying links or understanding truncated event titles in a calendar.
Future of computer use: towards expert-level performance
The current focus is on making computer use agents faster than the average human, with the next goal being 'superhuman' performance, matching or exceeding expert users. This will unlock more immediate and interactive experiences and lower the 'activation energy' for using automated software. However, challenges remain, including model limitations, inference speed, and latency. Delays in loading websites or slow responses can still be significant bottlenecks, requiring event-driven architectures and careful handling of timeouts and loading events.
API advancements for developers
The Agents API now includes computer use, enabling developers to build applications that interact with external websites and services universally. Key API updates include asynchronous function calling, allowing models to continue processing while tools execute, and steering during inference. The new Responses API has been entirely rewritten for maximum speed (TTFT and TTLB) with enhanced caching and pre-warming features. Developers are encouraged to use these new tools and provide feedback to shape future development. The core focus is on performance, reliability, and making the LLM platform the most efficient and dependable for building.
Decisions API: fast classification and parallel inference
The Decisions API is built on Luna's weights, focusing on high-speed classification and parallel inference rather than deep reasoning. It's designed for speed and structured output, making it suitable for tasks where quick categorization is needed. While not as sophisticated as Astra for complex code generation, it can be sufficient for certain computer use tasks. Internal prototypes have shown its potential when linked with real-time bidirectional APIs like GP Live, offering a smoother and more natural interaction for delegating tasks. The iterative deployment approach means users can expect further improvements as feedback is incorporated.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
●People Referenced
Common Questions
OpenAI DevDay announced Dots, a new personal assistant with advanced computer use features, and GPT-4.1 Soul, a faster and cheaper model for computer use. The Agents API now includes computer use capabilities, allowing developers to build on existing features.
Topics
Mentioned in this video
An Amazon Web Services product used in an analogy to describe the concept of building an AI cloud with customized AI versions of existing infrastructure.
A new personal assistant product featuring interesting computer use capabilities, each connected to a dedicated private Linux computer in the cloud.
A new API with parallel inference capabilities that is faster for simple tasks but less efficient for complex, long-term ones. It was inspired by Jeff's work and is being actively developed at OpenAI.
A tool used for computer use, capable of taking context from user computer activities and integrating it into ChatGPT. It also allows for writing JavaScript code for multi-action tasks.
Mentioned alongside Gemini as having released a bulletin with output styles similar to Jeff's approach.
A programming language used by computer users to write code that enables multi-action tasks and improves speed and capability in computer use.
A real-time, bidirectional API built on a front-end/back-end model, excellent for task delegation, which can be enhanced by Decisions API for tool calling.
Developers have extracted Inter from an intermediate layer, with speculation about its relevance to advanced AI model architectures.
Amazon Web Services, used as an analogy for building an AI cloud, suggesting the development of specialized AI versions of services like EC2 and S3.
A previous model or system referenced for cost comparison, with GPT-4.1 Soul being significantly cheaper for computer use tasks.
The technology GP Live utilizes, making it fast and efficient for task delegation.
The model the Decisions API is built upon, utilizing its weights and features for rapid, parallel processing and structured outputs.
An API that now includes computer use capabilities, allowing developers to build on the same features found in Codex and ChatGPT.
Mentioned as having shared a bulletin with output styles similar to Jeff's approach, indicating a trend in model development.
An Amazon Web Services product used in an analogy to describe the concept of building an AI cloud with customized AI versions of existing infrastructure.
A platform that benefits from the new computer use features, allowing users to quickly input context from their computer activities.
The operating system accessible within the virtual private cloud computer provided by each 'Dot', enabling the running of full desktop applications and web browsers.
A tool that computer use agents can leverage, along with screenshots and accessibility features, to perform various tasks.
The company where the speaker worked before OpenAI, focusing on building high-level products on top of core payment elements, a perspective applied to AI development.
A company co-founded by Ari while at Apple, which also worked on computer use, but with less capable models compared to current OpenAI offerings.
The company where Ari previously worked, co-founding a company called Sky, and which had specific policies that needed to be navigated for their work.
More from Latent Space
View all 258 summaries
95 minThe Future of Claude Code: Mods, Mutable Software, & Multiplayer Agents — Thariq Shihipar, Anthropic
83 minThe $10 Trillion Token Economy — Alex Atallah, OpenRouter & Anjney Midha, AMP
99 minRunway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis
92 min🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free