MANUS: A Truly Autonomous AI Agent?

Contents
  1. How China’s NEW Autonomous AI Agent Works
  2. How MANUS Works
  3. Integrating Multiple AI Models
  4. Task Management System
  5. Testing and Performance Results
  6. GAIA Test Results
  7. AI Agent Comparison
  8. Main MANUS Features
  9. ‘Manus’s Computer’ Interface
  10. Learning and Improvement System
  11. External Tool Support
  12. Current Problems and Limits
  13. System Stability Concerns
  14. AI Safety and Ethical Gaps
  15. Marketing Claims vs. Reality
  16. Final Analysis: MANUS Capabilities
  17. Main Findings
  18. What happened next (March 2025 to September 2026)

Updated 16 September 2026. This article was first published on 17 March 2025, eleven days after Manus launched as an invite-only beta. It has since been fact-checked against primary sources. The GAIA section now shows the actual scores Manus reported at launch instead of a table with no numbers; an earlier sentence that attributed those results to a “comparative study by OpenAI Deep Research” was wrong (the comparison was published by Manus itself, using figures OpenAI had reported for its own product) and has been corrected; and a dated timeline of what happened next, from the public launch to the announced and then blocked Meta acquisition, replaces the original outlook section.

MANUS is an AI agent designed to simplify complex tasks and deliver results autonomously. Unlike most AI systems that focus on single tasks, MANUS handles multi-step workflows by combining advanced models and task management. Here’s what you need to know:

  • What It Does: From financial analysis to travel planning, MANUS turns ideas into actionable outcomes.
  • How It Works: Integrates multiple AI models and breaks down tasks into manageable steps.
  • Key Features:
    • Data analysis with interactive dashboards.
    • Custom educational content creation.
    • Market research and supplier identification.
    • Personalized travel guides.
    • E-commerce performance insights.
  • Strengths: At launch Manus reported GAIA scores of 86.5%, 70.1% and 57.7% on levels 1 to 3, above the figures OpenAI had published for Deep Research. The numbers were self-reported and, as far as we can tell, never independently replicated.

However, documentation gaps on system stability, AI safety, and marketing claims raise questions about its real-world reliability. See the timeline at the end for what happened after launch.

Quick Comparison (Sample Outputs):

Task Type Example Outcome Output Format
Financial Analysis Tesla stock analysis with dashboards Interactive Reports
Educational Content Video explaining the momentum theorem Custom Presentations
Market Research Insights from YC W25 database for B2B companies Strategic Insights
Travel Planning Personalized itinerary Travel Guides

MANUS shows promise but needs more transparency and stability improvements to fully meet its potential.

How China’s NEW Autonomous AI Agent Works

How MANUS Works

MANUS

MANUS combines advanced AI models with a task management engine to handle a wide range of applications. Designed to simplify complex processes, it delivers results in both professional and personal contexts.

Integrating Multiple AI Models

MANUS uses a mix of AI models to tackle different challenges. By blending specialized tools, it supports tasks like financial analysis, educational content creation, and detailed research.

Here’s how it works in practice:

Task Type AI Model Used Example Outcome
Financial Analysis Stock Market Models Detailed Tesla stock analysis with interactive dashboards
Educational Content Learning Models Custom video presentations explaining the momentum theorem
Research Data Mining Models Insights from analyzing the YC W25 database to find qualifying B2B companies
E-commerce Analytics Models Performance analysis of Amazon stores with actionable insights

This multi-model setup allows MANUS to effectively handle complex tasks with precision.

Task Management System

The task management system in MANUS breaks down complex requests into smaller, manageable tasks. It analyzes the request, assigns the right AI models, coordinates the workflow, and ensures high-quality results. For example, when analyzing e-commerce operations, MANUS processes sales data, creates visual reports, provides tailored strategy recommendations, and compiles performance summaries.

This system is versatile enough to handle tasks like crafting personalized travel plans or conducting detailed market research on AI products across industries, making it a powerful tool for managing intricate projects.

Testing and Performance Results

The only quantitative evidence Manus offered at launch was a single benchmark result, which it published on its own website. Here is what it actually said, and where the comparison numbers come from.

GAIA Test Results

GAIA

GAIA (General AI Assistants) is a benchmark from researchers at Meta AI, Hugging Face and AutoGPT: 466 real-world questions that need web browsing, tool use and multi-step reasoning, graded in three difficulty levels. On its launch page on 6 March 2025, Manus published the following pass rates for its standard production configuration, alongside the numbers OpenAI had reported for Deep Research one month earlier and the previous state of the art cited by Manus.

GAIA level Manus (self-reported, March 2025) OpenAI Deep Research (reported by OpenAI, February 2025) Previous best (as cited by Manus)
Level 1 (basic) 86.5% 74.3% 67.9%
Level 2 (intermediate) 70.1% 69.1% 67.4%
Level 3 (complex) 57.7% 47.6% 42.3%

Sources: the Manus launch page at manus.im (the chart has since been removed from the site, but the figures were widely reported, for example by DataCamp), and OpenAI’s Deep Research announcement of 2 February 2025, which lists pass@1 scores of 74.29%, 69.06% and 47.6% for the same three levels.

AI Agent Comparison

Correction: an earlier version of this article said the comparison came from “a comparative study by OpenAI Deep Research”. That was wrong. OpenAI never evaluated Manus; the side-by-side chart was made by Manus, which placed its own scores next to the GAIA figures OpenAI had published for Deep Research. So the comparison is vendor-reported on one side and vendor-reported on the other, with no third party involved.

Three caveats apply. Manus did not publish its evaluation harness, per-question outputs or the number of runs, so the scores cannot be reproduced from the announcement. GAIA is a fixed public benchmark, and agents can be tuned against it. And a benchmark lead on browsing-and-tool tasks says little about the stability, safety and cost questions raised below. The honest reading in March 2025 was “promising but unverified”, and we have not seen an independent replication of the launch figures since.

Main MANUS Features

MANUS offers a standout experience among AI agents with its clear interface and ability to operate autonomously. Its features are designed to handle tasks seamlessly and efficiently, all while keeping users informed.

‘Manus’s Computer’ Interface

The ‘Manus’s Computer’ interface provides a detailed look into the AI’s decision-making process. Users can review every step of task execution through detailed replays, offering complete clarity.

This interface is versatile, supporting a wide range of applications:

Task Type Example Capability Output Format
Data Mining Patent Analysis Comparative Reports
Content Creation Technical Documentation Interactive Guides
Market Research Competitor Analysis Strategic Insights
Project Management Resource Optimization Progress Dashboards

Learning and Improvement System

MANUS is designed to continuously grow and refine its abilities. It thrives in:

  • Identifying patterns in tasks
  • Enhancing efficiency over time
  • Solving challenges in a flexible manner
  • Making decisions that align with the context

External Tool Support

Beyond its built-in features, MANUS integrates smoothly with external platforms. This allows it to:

  • Turn complex datasets into clear, actionable visualizations
  • Organize documents effectively, producing comparison tables and detailed guides
  • Search across multiple databases to compile structured data across various formats

Current Problems and Limits

While MANUS showcases strong potential, its documentation falls short in addressing key areas like system stability, safety measures, and how well its marketing claims align with actual performance. Despite earlier discussions of its features and test outcomes, these gaps leave several critical questions unanswered.

System Stability Concerns

The documentation doesn’t provide enough information about potential stability issues when integrating multiple AI models and external tools. It’s unclear how MANUS maintains consistent performance under different workloads, especially given the wide range of tasks it aims to handle.

AI Safety and Ethical Gaps

Details about AI safety protocols and ethical considerations are noticeably absent. For instance, the documentation doesn’t explain how transparent its decision-making processes are or what restrictions are in place for its autonomous actions. This lack of clarity makes it harder to evaluate how safely and responsibly MANUS operates in real-world scenarios.

Marketing Claims vs. Reality

MANUS’s promotional materials highlight its ability to handle a variety of tasks. While this paints an optimistic picture, independent testing is crucial – especially for tasks requiring complex understanding or problem-solving. The disconnect between the marketing promises and the lack of concrete documentation underscores the importance of validating its performance in practical settings.

Final Analysis: MANUS Capabilities

Main Findings

MANUS stands out for its ability to turn AI-driven insights into real-world actions, setting itself apart from many AI systems that excel in analysis but fall short in execution. Its strong benchmark performance and built-in task management make it a practical tool for achieving measurable outcomes across various tasks.

What happened next (March 2025 to September 2026)

The original version of this section speculated about future updates. Here is what actually happened, with sources.

  • 6 March 2025: Manus launches as an invite-only beta from Butterfly Effect, the Beijing-founded team behind the Monica browser assistant. More than two million people join the waitlist in the first week.
  • May 2025: the waitlist is removed and anyone can sign up with free daily task credits, shortly after a $75 million round led by Benchmark (South China Morning Post). The company later moves its headquarters to Singapore.
  • 16 October 2025: Manus 1.5 ships; 15 December 2025: Manus 1.6 adds a higher-performance mode, mobile development and a design view (Manus blog). On 17 December the company reports $100 million in annual recurring revenue.
  • 29 December 2025: Meta announces it will acquire Manus, in a deal reported at around $2 billion.
  • April 2026: China’s National Development and Reform Commission prohibits the transaction on national-security grounds and orders the parties to withdraw it (France 24).
  • 1 September 2026: Manus formally resumes independent operations as an “independent agent lab” led by its founding team. Users whose data was deleted during the separation are offered a restoration portal.

So the product survived, grew fast and became a large acquisition target within a year, then had that exit reversed by regulators. The open questions from March 2025 (stability under load, safety controls, and how far the marketing runs ahead of the evidence) still apply, and the launch benchmark remains the only hard number the company has published for the agent’s task performance.

Related Blog Posts


Previous
Next

← All writing