Key Moments

Inside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj

Latent Space PodcastLatent Space Podcast
Science & Technology5 min read41 min video
Sep 30, 2026|1,697 views|21|5
Save to Pod
TL;DR

OpenAI's new computer use agents are now faster than humans and cost 1/7th of previous models, enabling complex tasks like meal planning in minutes instead of hours, but questions remain about their ability to outperform expert human users.

Key Insights

1

Computer use agents can now perform tasks 7x faster than before, with new models like GPT-4.1 Soul costing 1/5th of Astra and 1/7th for computer use specifically.

2

The new Agents API allows developers to build on the same computer use capabilities previously only available in Codex and ChatGPT.

3

Dots, a new personal assistant product, offers each user a dedicated virtual Linux computer in the cloud, capable of running full desktop applications and web browsers.

4

The Decisions API, while not dependent on reasoning, performs inference in parallel and uses a smaller, faster model for tasks where complex, long-term reasoning isn't critical.

5

OpenAI's computer use now writes JavaScript code to execute multiple actions simultaneously, significantly improving speed and capability, and utilizes multimodal inputs like screenshots and accessibility features.

6

The new Responses API has been rewritten for maximum speed in Time To First Byte (TTFT) and Time To Last Byte (TTLB), with improved caching and pre-warming features for continuous agent operation.

Computer use agents achieve superhuman speed and cost efficiency

Computer use agents have advanced to the point where they are now faster than the average human for most tasks. The next frontier is achieving 'superhuman' performance, matching or exceeding expert human computer users. This advancement promises more interactive and immediate product experiences. This was not a certainty just a few months ago, but OpenAI has rapidly iterated, with internal teams embracing and developing the idea. The new GPT-4.1 Soul model is particularly notable for computer use, costing only one-fifth of Astra and a seventh for specific computer use tasks. The new Decisions API, while distinct, performs inference in parallel using a smaller model, making it fast but less suited for complex, long-term tasks, highlighting ongoing research into integrating different approaches.

New products and APIs empower developers and users

OpenAI DevDay saw several key announcements. Dots, a new personal assistant, provides each user with their own private virtual Linux computer in the cloud, enabling the execution of full desktop applications and web browsers. This is a significant shift from previous cloud browser or personal computer access. The Agents API now integrates computer use capabilities, allowing developers to build upon the same functionality present in Codex and ChatGPT. Demonstrations included App Shots, where context from user actions can be quickly input into AI models, and enhanced existing computer use features on Mac.

The power of dedicated cloud computing for AI agents

A core innovation is that each 'Dot' now has access to a dedicated virtual Linux computer in the cloud. This is a departure from traditional cloud access methods, offering a full computing environment. This capability is crucial because most software is designed for humans, and now AI agents can leverage these same programs. Users can delegate any task they perform on a computer to a Dot, from booking flights and shopping to more complex, time-consuming activities. An example cited was a user spending two hours customizing a meal delivery order, a task an agent completed in just 15 minutes, saving significant time and effort.

Evolution of computer use through advanced techniques

Computer use has evolved significantly, moving beyond basic task initiation to sophisticated error correction and iterative attempts. A key development is that computer use agents now often write code, such as JavaScript, to perform multiple actions simultaneously, drastically improving speed and capability. The interaction layer has also become more sophisticated, utilizing multimodal inputs like screenshots, accessibility features, and tools like Playwright. This allows models to 'see' entire pages or applications and execute complex sequences of actions, rather than just single steps.

Enhanced computer use through multimodal inputs and accessibility

The integration of multimodal inputs has been a major breakthrough. Previously, agents would struggle with tasks requiring scrolling, taking multiple screenshots, and re-evaluating. With improved accessibility features and direct DOM access, models can now perceive entire pages or applications at once. They can write code to perform multi-step actions efficiently. This also means that technologies designed for users with disabilities, like screen readers and accessibility representations, are proving highly beneficial for AI agents, providing rich context for tasks like identifying links or understanding truncated event titles in a calendar.

Future of computer use: towards expert-level performance

The current focus is on making computer use agents faster than the average human, with the next goal being 'superhuman' performance, matching or exceeding expert users. This will unlock more immediate and interactive experiences and lower the 'activation energy' for using automated software. However, challenges remain, including model limitations, inference speed, and latency. Delays in loading websites or slow responses can still be significant bottlenecks, requiring event-driven architectures and careful handling of timeouts and loading events.

API advancements for developers

The Agents API now includes computer use, enabling developers to build applications that interact with external websites and services universally. Key API updates include asynchronous function calling, allowing models to continue processing while tools execute, and steering during inference. The new Responses API has been entirely rewritten for maximum speed (TTFT and TTLB) with enhanced caching and pre-warming features. Developers are encouraged to use these new tools and provide feedback to shape future development. The core focus is on performance, reliability, and making the LLM platform the most efficient and dependable for building.

Decisions API: fast classification and parallel inference

The Decisions API is built on Luna's weights, focusing on high-speed classification and parallel inference rather than deep reasoning. It's designed for speed and structured output, making it suitable for tasks where quick categorization is needed. While not as sophisticated as Astra for complex code generation, it can be sufficient for certain computer use tasks. Internal prototypes have shown its potential when linked with real-time bidirectional APIs like GP Live, offering a smoother and more natural interaction for delegating tasks. The iterative deployment approach means users can expect further improvements as feedback is incorporated.

Common Questions

OpenAI DevDay announced Dots, a new personal assistant with advanced computer use features, and GPT-4.1 Soul, a faster and cheaper model for computer use. The Agents API now includes computer use capabilities, allowing developers to build on existing features.

Topics

Mentioned in this video

Software & Apps
EC2

An Amazon Web Services product used in an analogy to describe the concept of building an AI cloud with customized AI versions of existing infrastructure.

Dots

A new personal assistant product featuring interesting computer use capabilities, each connected to a dedicated private Linux computer in the cloud.

Decisions API

A new API with parallel inference capabilities that is faster for simple tasks but less efficient for complex, long-term ones. It was inspired by Jeff's work and is being actively developed at OpenAI.

Codex

A tool used for computer use, capable of taking context from user computer activities and integrating it into ChatGPT. It also allows for writing JavaScript code for multi-action tasks.

Gemma

Mentioned alongside Gemini as having released a bulletin with output styles similar to Jeff's approach.

JavaScript

A programming language used by computer users to write code that enables multi-action tasks and improves speed and capability in computer use.

GP Live

A real-time, bidirectional API built on a front-end/back-end model, excellent for task delegation, which can be enhanced by Decisions API for tool calling.

Inter

Developers have extracted Inter from an intermediate layer, with speculation about its relevance to advanced AI model architectures.

AWS

Amazon Web Services, used as an analogy for building an AI cloud, suggesting the development of specialized AI versions of services like EC2 and S3.

Astra

A previous model or system referenced for cost comparison, with GPT-4.1 Soul being significantly cheaper for computer use tasks.

Docker

The technology GP Live utilizes, making it fast and efficient for task delegation.

Luna

The model the Decisions API is built upon, utilizing its weights and features for rapid, parallel processing and structured outputs.

Agents API

An API that now includes computer use capabilities, allowing developers to build on the same features found in Codex and ChatGPT.

Gemini

Mentioned as having shared a bulletin with output styles similar to Jeff's approach, indicating a trend in model development.

S3

An Amazon Web Services product used in an analogy to describe the concept of building an AI cloud with customized AI versions of existing infrastructure.

ChatGPT

A platform that benefits from the new computer use features, allowing users to quickly input context from their computer activities.

Linux

The operating system accessible within the virtual private cloud computer provided by each 'Dot', enabling the running of full desktop applications and web browsers.

Playwright

A tool that computer use agents can leverage, along with screenshots and accessibility features, to perform various tasks.

More from Latent Space

View all 258 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free