From Chatbots to Getting Things Done: What Is Changing in Enterprise AI
2025 was widely called the year of the AI agent. Looking back from early 2026: a dated timeline of the industry's turn from answering questions to completing tasks — Computer Use, Mariner, Operator, Manus, ChatGPT's agent mode — how to read the benchmarks calmly, and the posture businesses should take now.
Key takeaway
From October 2024 to July 2025, AI moved from answering to acting: Anthropic's Computer Use, Google's Mariner, OpenAI's Operator, Manus and ChatGPT's agent mode. OSWorld scores — 38.1% versus 72.4% for humans — say agents can work but not unsupervised, so the sensible pattern is AI drafts, people confirm.

Rewind to before late 2024: the default form of AI at work was a chat window. You asked, it answered — and the copying, pasting, opening systems and submitting forms stayed your job. Through 2025, "let the AI finish the task itself" went from lab talk to product race. It is now January 2026, and "the year of the agent" has been said for a year; time to look back at what actually happened, and how far things actually got.
The timeline: from moving a cursor to joining the main product
October 2024. Anthropic released Computer Use: Claude 3.5 Sonnet, in public beta, gained the ability to operate a computer — reading screenshots, moving the cursor, clicking, typing, using graphical interfaces the way a person does. Crude still, but the direction was unmistakable: for the first time a mainstream model was productised to act, not just to answer.
December 2024. Google announced Project Mariner, a more focused bet: performing web actions for the user inside Chrome. The browser is the doorway to most business systems, and the choice of venue said plenty about where the industry expected agents to land first.
January 23, 2025. OpenAI launched Operator as a research preview. Its CUA (Computer-Using Agent) model combines GPT-4o's vision with reinforcement learning and drives a cloud browser to book tables, fill forms and place orders for the user. OpenAI published benchmarks alongside: on OSWorld, which simulates real computer operation, a success rate of 38.1% against 72.4% for humans; on the web-task benchmark WebVoyager, 87%. More on those numbers below.
March 6, 2025. Manus, a general-purpose agent from the Chinese team behind Monica (Butterfly Effect), launched invitation-only: hand it a task and it decomposes the steps, researches, operates tools and delivers a result. Invitation codes were soon trading second-hand for anywhere from a few hundred to tens of thousands of yuan — the company clarified it had never sold codes — and by May 2025 registration opened to everyone.
July 17, 2025. OpenAI folded Operator's capabilities into ChatGPT's agent mode and retired the standalone site. Often read as a product-line reshuffle, it is better read as a signal: the ability to execute stopped being a separate product and became a default capability of the mainstream assistant.

How to read 38.1% against 72.4%
Two misreadings of the OSWorld numbers are both common. The optimistic one looks only at the trend and skips the absolute value: under-40% success means that left unsupervised, roughly two of every three tasks need a human to clean up — "unattended" was simply not true in 2025, which is why every mainstream product kept confirmation and takeover mechanisms. The pessimistic one treats 38.1% as a ceiling: it is a snapshot on one benchmark, and WebVoyager's 87% sits right beside it — the more structured the task and the clearer its boundaries, the more usable the agent. The honest conclusion lies between: it can work, if you choose its work carefully.
What the invite-code frenzy actually measured
The Manus episode was the most sociological moment on the timeline. Access to an invitation-only product trading at up to tens of thousands of yuan measured how badly the market wants "AI that does my work"; the cooling after May's open registration measured how far expectation had outrun capability. For decision-makers the lesson is direct: when you hear "agent," do not reach for the budget — return to the task. Which of your processes can it complete reliably, and who picks up the pieces when it fails?
What this means for companies: enter, but with the right posture
Converted into company action, 2025 says: start learning, pilot on a small scale, and keep "AI drafts, human confirms" as the working form — let the agent do the legwork and first drafts (research, consolidation, form-filling, reply drafting) while people keep the irreversible acts: submitting, sending, paying. If the difference from an ordinary chatbot still feels fuzzy, start with what an AI agent is; and for many businesses the honest starting point is not an agent at all — fixed processes run more reliably as workflows, a trade-off worked through in AI workflow or AI agent.
Three questions for any "agent product"
Products wearing the agent label will keep multiplying, and the demos are uniformly impressive. Three questions cut deeper than any demo:
- What tools does it actually connect to? Your real systems — spreadsheets, CRM, ticketing, email — or only a demo sandbox;
- How are permissions managed? Which data it may touch, which actions need human approval, and whether authority can be narrowed per person and per scenario;
- What happens when it goes wrong? Complete action logs, reversible operations, a human able to take over at any moment — the cost of an error decides what may be entrusted.
What to expect from 2026
Seen whole, the timeline from Computer Use to agent mode records an industry confirming a direction in about a year: AI's next stop is execution, not better conversation. A confirmed direction is not delivered capability — the gap between 38.1% and 72.4% is the hole 2026 has to keep filling. For companies this is a rare kind of window: the technology is not yet so mature that abstaining means losing, and already real enough to be worth hands-on trials. Accumulating experience and judgment inside low-risk processes beats waiting for the perfect product.