
Manus built a software business without training its own foundation model. According to ifanr, the product's ARR had exceeded $100 million by December last year by providing a cloud-based virtual machine to research information and write code. Within 24 hours of Manus becoming popular in March 2025, online observers alleged that it had no self-developed large model and relied on Claude's API, woshipm reported.
The engine lived in the harness rather than the weights. External scaffolding turned text generation into reliable execution.
In March, after leaving Qwen, Lin Junyang published an essay titled "From Reasoning-Oriented Thinking to Agent-Oriented Thinking." As woshipm reported, Lin defines Harness Engineering as a system surrounding a model that includes tools, memory, environments, permissions, verification mechanisms, and multi-agent coordination patterns. Manus implemented this architecture across three internal roles. A planner broke requests into executable steps, while an executor called tools and operated browsers. A verifier checked results and sent errors back to the planner. This orchestration succeeds in software sandboxes, but it breaks down when a domain cannot be preserved in language.
How Manus turned execution harnesses into commercial defensibility
In March, after leaving Qwen, Lin Junyang published an essay titled "From Reasoning-Oriented Thinking to Agent-Oriented Thinking." According to woshipm, Lin defined harness engineering as an operational system surrounding a model. This framework encompasses tools, memory, environments, permissions, verification mechanisms, and multi-agent coordination patterns. Models function as modular engines within that broader harness.
- March 2025Manus launched as a general AI agent
- JulyAndrew Ng open-sourced OpenWorker
- At the end of July and the beginning of AugustTencent consolidated Agent products into WorkBuddy
- August 3Alibaba began public testing Qianwen Office
- AugustDeepSeek open-sourced its Harness framework
- August 2026Manus announced resumption of independent operations
Rented models can support commercial software. Within 24 hours of Manus gaining attention in March 2025, critics alleged it relied on Claude's API, woshipm reported. Yet ifanr noted that Manus's ARR exceeded $100 million by December last year.
Manus establishes defensibility by coordinating models through specialized internal roles. As woshipm explains, a planner first breaks user requests into executable steps. An executor runs code and calls tools. A verifier checks results and sends errors back to the planner. Cost control depends on separating reasoning intensity. Manus routes complex logical reasoning tasks to Claude and passes inexpensive repetitive tasks to fine-tuned small Qwen models. Through this multi-model routing, woshipm notes that Manus reduced the average cost of a single task to about $2, roughly 40% less than using one large model alone.
Execution also requires isolated state management. According to woshipm, each Manus task operates within an isolated Linux virtual environment featuring a terminal and a browser. Task-level memory resides in sandbox file systems, while user-level memory is preserved in a vector database. ifanr notes that these cloud-based virtual machines let the agent write code and research information. The environment can also create PowerPoint presentations. The harness delivers finished artifacts. As woshipm highlights, Manus delivers completed outputs like formatted reports with data rather than stopping at conversational guidance.
Three approaches to making models useful beyond chat
| Manus execution system | Pragmatik Labs' agent architecture | Scinetics' AI4S representation approach | |
|---|---|---|---|
| Primary setting | Cloud-based virtual machine for research, coding and PowerPoint creation | Digital knowledge work and physical real-world environments | Life sciences initially, with later expansion to materials and engineering |
| Core system design | Planner, executor and verifier roles form a feedback loop | Harness Engineering surrounds a model with tools, memory, environments, permissions, verification and multi-agent coordination | A multimodal, long-horizon reasoning foundation model for AI4S |
| Tool and execution layer | Executor calls tools, runs code and operates browsers in an isolated Linux virtual environment | Agents share tool use and action-coordination capabilities | Scientific questions are translated into structured instructions for tools to execute |
| Memory and verification | Task memory is stored in sandbox file systems; user memory is stored in a vector database; verifier returns errors to the planner | Memory and verification mechanisms are named as Harness Engineering components | Not covered |
| Model strategy | Routes complex logical reasoning to Claude and inexpensive repetitive tasks to fine-tuned small Qwen models | Not covered | Uses discrete tokenization schemes for scientific modalities |
| Key constraint or failure mode | Leading models still make errors in global constraints and time budgets on long-horizon tasks | Not covered | Detailed 3D structure and subtle cross-modal relationships can be lost in language translation |
| Representation approach | Textual requests are executed through tools and virtual computers | Not covered | Science Token is proposed as a unified representation for small molecules, nucleic acids and protein structures |
| Output goal | Completed outputs such as formatted reports with data, rather than textual recommendations alone | Digital agents for research, programming and analysis; physical agents for perception, decision-making and execution | Scientific reasoning across tasks including protein structure prediction and DNA mutation prediction |
| Pricing information | Average single-task cost reduced to about $2 | Not disclosed in sources | Science Token pricing would consider scientific value, computing power and electricity; no mature industry-wide charging standard |
Why language orchestration hits a wall in scientific workflows
According to woshipm, Manus splits execution across three internal roles. A planner divides requests into executable steps, while an executor calls tools and operates web browsers. A verifier inspects the outputs and returns errors directly to the planner. Backed by a cloud-based virtual machine, as ifanr observed, this harness writes code and compiles PowerPoint presentations to answer user requests. It delivers completed outputs like formatted reports with data instead of plain text suggestions.
Scientific discovery resists this text-and-browser loop. When tasks hinge on physical geometry rather than digital files, translating spatial reality into sentences destroys essential data.
Zhang Zaixi encountered that boundary while developing the scientific agent Stella before founding Scinetics. Stella reached more than 1,000 scientific-research users after launching last year. More than 70% of those users came from universities and top laboratories, including Stanford and Princeton. The system converted scientific inquiries into natural language for the model before turning them back into structured instructions for tools to execute. That process failed at the seams. Zhang noted that the translation discarded detailed 3D structural information along with subtle relationships between modalities.
Orchestration cannot fix missing coordinates. Because language creates a representation ceiling, Scinetics built a multimodal, long-horizon reasoning foundation model for AI4S. The company introduced Science Token as a unified representation unit for scientific data. This format preserves spatial structures directly, covering small molecules and nucleic acids alongside protein structures.
What long-horizon benchmarks reveal about agent cost and recovery
Standard language evaluations reward isolated answers, but execution harnesses struggle once tasks extend across dozens of sequential operations. According to woshipm, OSWorld 2.0 expanded computer operation evaluation to 108 real long-horizon workflows with a median human completion time of 1.6 hours. The best current model in the main evaluation fully completes only about 20% of those tasks. The failure pattern is structural. A DeepPlanning experiment involving Lin Junyang found that even leading frontier models frequently make errors in global constraints and time budgets on long-horizon tasks such as multi-day travel planning and complex procurement.
Long execution chains turn minor planning mistakes into compounding dead ends. When an agent mismanages its time budget or violates an early constraint, every subsequent action burns compute without bringing the goal closer.
Absorbing these failure loops demands strict operational economics. Data cited by ifanr and woshipm shows Manus cumulatively created 80 million virtual computers and processed 147 trillion tokens. To keep that overhead viable, Manus reduced the average cost of a single task to about $2 through model routing, roughly 40% less than using one large model alone.

Yet pricing protracted reasoning remains an open problem. Raw compute is an incomplete yardstick. Zhang Zaixi explained that Science Token pricing would consider scientific value in addition to basic costs such as computing power and electricity.
A diagnostic rubric for workflow bottlenecks versus representation ceilings
To diagnose why an agent stumbles, practitioners must separate operational breakdowns from vocabulary limits. Lin Junyang defines harness engineering as a system surrounding a model that includes tools, memory, environments, permissions, verification mechanisms, and multi-agent coordination patterns. According to woshipm, Pragmatik Labs points to four shared capabilities across agents: reasoning, tool use, learning from feedback, and coordinating actions. In Manus, woshipm notes that each task runs in an isolated Linux virtual environment with a terminal and file system alongside a browser.
External compliance rules define where this harness can operate. As ifanr reported, Manus announced that data generated after December 29, 2025, the day Meta acquired Manus, required corresponding processing for regulatory compliance. Users had to complete manual backups before 8 a.m. Singapore time on August 23 and re-upload data for restoration on August 25. Manus stated that this adjustment was unrelated to a hacker attack or data leak, attributing it to data migration and compliance arrangements following its return to independent operations.
When an execution harness collapses, the fault lies in sandbox limits or coordination rules. When an agent misinterprets domain-specific logic, the failure is representational.
Representation ceilings stem from the translation layer. Zhang Zaixi explained that scientific agents translate scientific questions into natural language for a model and then translate them back into structured instructions for tools to execute. During that translation, detailed 3D structural information and subtle relationships between modalities can be lost. No harness fix recovers data that words drop. To avoid that loss, Zhang Zaixi noted that Scinetics developed discrete tokenization schemes for omics data, protein structures, molecular chemical formulas, and nucleic-acid sequences.
Choose orchestration for execution failures; change representations for scientific-information failures
- A task requires research, coding, browser operation, file handling and a finished deliverable rather than a chat response. Use an execution-system approach: Manus provides a cloud-based virtual machine for research, code and PowerPoint creation, and is designed to deliver completed outputs such as formatted reports with data. Its reported planner, executor and verifier roles offer a concrete pattern for decomposing work, using tools and checking results.
- A workflow is long-horizon and must satisfy global constraints, time budgets or interdependent steps. Build explicit planning and verification loops rather than assuming a frontier model can manage the workflow unaided. A DeepPlanning experiment found frequent errors in global constraints and time budgets, while the best current model in the main OSWorld 2.0 evaluation could fully complete only about 20% of tasks.
- The workflow mixes expensive reasoning with repetitive, lower-stakes operations. Route work across models rather than sending every step to one large model. Manus reportedly routes complex logical-reasoning tasks to Claude and inexpensive repetitive tasks to fine-tuned small Qwen models, reducing average single-task cost to about $2, roughly 40% less than one large model alone.
- An agent must retain working artifacts during a task and preferences or context across tasks. Separate task memory from user memory and isolate execution. Manus reportedly keeps task-level memory in sandbox file systems and user-level memory in a vector database; each task runs in an isolated Linux virtual environment with a browser, terminal and file system.
- The task depends on detailed molecular, structural, omics or other scientific modalities whose meaning may be degraded by conversion into language. Do not treat this as merely an agent-orchestration problem. Zhang Zaixi said that translating scientific questions into natural language and back into tool instructions can lose detailed 3D structural information and subtle cross-modal relationships. Scinetics' proposed alternative is Science Token, with discrete tokenization schemes for omics data, protein structures, molecular chemical formulas and nucleic-acid sequences.
Connecting domain-native tokens to general execution sandboxes
Enterprise execution platforms are consolidating around general workspaces. At the end of July and the beginning of August, Tencent consolidated several Agent products into WorkBuddy as a strategic entry point, according to woshipm. Alibaba merged QoderWork, Wukong and MuleRun into Qianwen Office (千问办公), which began public testing on August 3 under the direct management of DingTalk's new CEO. These platforms coordinate interfaces and permissions across office workflows.
General harnesses excel at running tools. Yet when a workflow hinges on specialized scientific structures, standard text representations fall short.
To bridge that gap, Scinetics builds a multimodal, long-horizon reasoning foundation model for AI4S. Zhang Zaixi said Scinetics developed discrete tokenization schemes for omics data, protein structures, molecular chemical formulas and nucleic-acid sequences. Individual task tests showed positive results. These include protein structure prediction, protein function annotation, RNA secondary-structure prediction and DNA mutation prediction.
Connecting native representations to execution environments defines the next layer of agent architecture. According to woshipm, Pragmatik Labs operates two business lines: digital agents for knowledge work and physical agents for real-world environments. Its digital agents target research and programming alongside analysis. The woshipm author argues that Pragmatik Labs is building a cross-environment agent foundation architecture before applying capabilities to software and robotics scenarios. Scinetics focuses initially on life sciences before expanding into materials and engineering. Zhang Zaixi said Scinetics aims to bill through Science Token units, building a foundation model company comparable to DeepSeek and Anthropic.
According to woshipm, Manus isolated its architecture by storing task-level memory in sandbox file systems and user-level memory in a vector database. Across 147 trillion tokens, the system delivered formatted reports with data rather than conversational recommendations. Andrew Ng applied a similar boundary in July when open-sourcing OpenWorker, prioritizing finished products over raw chat records.
Inspect the harness before blaming the model.
Teams can benchmark their setups against emerging implementations. DeepSeek open-sourced its Harness framework in August, setting a GitHub star-growth record on its launch day. In enterprise deployments, woshipm recorded Tencent consolidating services into WorkBuddy across late July and early August, reaching 20.97 million monthly visits with daily active users three to four times those of the second-ranked product.
For readers outside China
- Availability: Manus shifted its operational focus to Singapore and rebuilt its team, business and service system around global markets. However, the sources do not say which countries can use it, whether it is currently open to new users, or whether Scinetics or Pragmatik Labs products are available outside China.
- Pricing: No public end-user pricing is disclosed in sources. Manus reportedly reduced its average cost of a single task to about $2, but that is an internal task-cost claim rather than a published customer price. Scinetics says Science Token pricing would consider scientific value as well as computing power and electricity, but no price or billing schedule is disclosed.
- Closest Western equivalents: Claude, which Manus reportedly uses for complex logical-reasoning tasks; Anthropic, which Scinetics says it aims to be comparable to as an AI4S foundation-model company
- Data residency: The source material does not specify where Manus, Scinetics or Pragmatik Labs store customer data. Manus said that data generated after December 29, 2025 required corresponding processing for regulatory compliance, and described a data migration and compliance arrangement after returning to independent operations. That does not establish a data-residency location or a customer-data policy.
Sources
- woshipm 林俊旸官宣Pragmatik Labs:不做大模型续集,去造横跨「数字与物理」的下一代Agent https://woshipm.com/share/6446523.html
- ifanr Manus 突然宣布独立运营,一夜回到创业状态 https://ifanr.com/1674998
- 36kr 以小博大,一家AI4S水下公司的野心 https://36kr.com/p/3985879828478981
- woshipm 被骂"套壳"的Manus,教了产品经理五件事 https://woshipm.com/ai/6453824.html
The evidence: 25 facts from 4 Chinese articles
Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.
36kr以小博大,一家AI4S水下公司的野心
- In May, 28-year-old Zhang Zaixi ended his postdoctoral research at Princeton and returned to China.
- Zhang Zaixi became an assistant professor at the Hong Kong University of Science and Technology and founded the AI for Science company Scinetics.
- Scinetics recently completed a financing round of nearly 50 million yuan led by Innofund, with Yijing Capital, Xiaomiao Langcheng and Linge Ventures participating.
- Before founding Scinetics, Zhang Zaixi and his team developed the scientific agent Stella.
- During his PhD studies, Zhang Zaixi developed the molecular screening model MGSSL and the drug-molecule generation model FLAG.
ifanrManus 突然宣布独立运营,一夜回到创业状态
- Manus announced that it will resume operations as an independent company.
- Manus said that data generated after December 29, 2025, the day Meta acquired Manus, requires corresponding processing for regulatory compliance.
- Manus asked users to complete manual backups before 8 a.m. Singapore time on August 23 and re-upload data for restoration on August 25.
- Manus said its data adjustment is unrelated to a hacker attack or data leak and is primarily due to data migration and compliance arrangements following its return to independent operations.
- Manus said affected users will receive compensation, including free service during the more than ten-day period and a "return gift package" after the service relaunches.
- Manus founder Xiao Hong is a 1992 graduate of Huazhong University of Science and Technology.
- Manus provides a cloud-based virtual machine to research information, write code, and create PowerPoint presentations in response to user requests.
- Manus shifted its operational focus to Singapore and rebuilt its team, business, and service system around global markets.
woshipm林俊旸官宣Pragmatik Labs:不做大模型续集,去造横跨「数字与物理」的下一代Agent
- On August 12, former Alibaba Qwen technical lead Lin Junyang announced the founding of Pragmatik Labs in Shanghai.
- Pragmatik Labs' Chinese legal entity is Yuyong Technology (语用科技), and its internal abbreviation is p7k.
- Gaorong Capital and Sequoia China jointly led Pragmatik Labs' funding round, with Tencent and the Shanghai Future Industry Fund participating in support.
- Pragmatik Labs' website lists two business lines: digital agents for knowledge work and physical agents for real-world environments.
- In May, The Information reported that Lin Junyang was raising hundreds of millions of dollars for a new AI lab at a potential post-money valuation of $2 billion.
- In June, multiple media outlets citing informed sources reported that Gaorong Capital and Sequoia China each invested $100 million in Pragmatik Labs and Tencent invested $20 million, for a combined total of at least $220 million.
- Lin Junyang joined Alibaba's DAMO Academy after graduating with a master's degree from Peking University in 2019.
- Lin Junyang worked on multimodal projects including M6 and OFA at Alibaba.
- At the end of 2022, Lin Junyang became one of Qwen's technical leads.
- In March, after leaving Qwen, Lin Junyang published an essay titled "From Reasoning-Oriented Thinking to Agent-Oriented Thinking."
- OSWorld 2.0 expanded computer operation evaluation to 108 real long-horizon workflows with a median human completion time of 1.6 hours.
woshipm被骂”套壳”的Manus,教了产品经理五件事
- Manus announced the resumption of independent operations in August 2026.