Claude memory turns context into workflow control because an agent is only as useful as the past it can safely reuse and the tools it is allowed to touch. woshipm traces the shift from OpenAI's ChatGPT research preview on November 30, 2022, to function calling in June 2023. Gemini 1.5 brought million-token context into product discussions in 2024. The pattern is clear: chat became tool use, and tool use made context a control surface.
The next turn made that surface harder to manage.
Anthropic publicly tested Computer Use in October 2024 and open-sourced MCP in November 2024, according to woshipm. OpenAI then released the Responses API and Agents SDK in March 2025. These are not just bigger prompt boxes. They are ways to decide what an AI can remember or fetch. They also define when it can click or call. Human approval becomes part of that boundary. The real workflow skill is no longer prompt polish. It is governing context before context governs the work.
From prompt craft to context governance
The workflow problem changed after the ChatGPT research preview on November 30, 2022. At first, the practical craft looked like prompt writing: give the model better instructions, choose a better model, then ask again. That was a narrow loop. Once OpenAI added function calling to its API in June 2023, the model could stop being only a text generator and start becoming a dispatcher for tools.
- November 30, 2022OpenAI released the ChatGPT research preview
- June 2023OpenAI added function calling capability to its API
- November 2024Anthropic open-sourced MCP
- March 2025OpenAI released the Responses API and Agents SDK
- October 2025Anthropic launched Agent Skills
- 2026OpenAI used Skills for ChatGPT, Codex, and API
Long context widened the same issue. Gemini 1.5 pushed product discussions to the million-token level in 2024, according to woshipm, which made it possible to feed a system far more prior material. But capacity is not governance. A larger window can hold more documents; it can also hold prior chats, policies, or examples. It does not decide which material should shape the next action.
That is where context becomes workflow control.
The agent stack made the shift harder to ignore. Anthropic publicly tested Computer Use in October 2024, then open-sourced MCP in November 2024. OpenAI released the Responses API plus the Agents SDK in March 2025. Anthropic launched Agent Skills in October 2025. OpenAI used Skills for ChatGPT, Codex, as well as API in 2026. The woshipm author argues that AI is moving from content generation toward systems that understand goals, connect resources, then execute tasks.
Once a system can call tools, operate a computer, connect through MCP, or load a Skill, memory is no longer a pleasant personalization feature. It becomes part of the control plane. The question is not simply what the model knows. It is what the workflow permits the model to recall, when it may use that recall, and when old context must be excluded.
Woshipm describes continual learning in product contexts as parameter learning. It also describes external memory as a separate layer. Process learning is the third form. External memory covers chat memory, project memory, RAG, knowledge bases, vector databases, or graph memory. Process learning is more procedural: an Agent can record project decisions, failure reasons, customer preferences, or acceptance criteria and reuse them later. A usable system, woshipm says, usually needs capture and filtering first. It then needs consolidation. Retrieval has to be controlled, and forgetting is also required. That sequence is the new operating skill.
Which memories deserve to survive
The useful lesson from consumer memory features is selective survival. Earlier ChatGPT memory, as woshipm describes it, looked closer to keeping a few notes. After 2025, OpenAI introduced background organization that can refer to chat history. Its FAQ says memory is also a continuously updated synthesis, not only the saved items a user can view in a list.
For office work, that shift changes the question from "what should the assistant remember?" to "what should be allowed to keep influencing later work?" Stable preferences qualify, as do project constraints. Acceptance criteria and failure reasons qualify. Recurring tasks qualify as well.
Microsoft 365 Copilot points in this direction by putting memory into office scenarios such as communication style, work goals, preferred topics, plus recurring tasks.
A memory that cannot be inspected is closer to hidden policy than helpful context.
That is why visible correction matters. OpenAI emphasizes memory summaries that users can review. Users can add to those summaries, and update them as well. Claude's help documentation says Claude organizes memory into categorized entries after memory is enabled. During conversations, Claude reads memory; it writes memory too, and it also updates memory. Claude also keeps separate memory spaces plus project summaries for each project, which turns an ordinary workplace concern into product behavior: the assistant should not let one client's assumptions leak into another client's plan.

Deletion is part of the workflow, not a privacy footnote. Woshipm notes that Claude can pause or reset memory, separates project memory from non-project chats, plus supports importing and exporting memory across AI services. Gemini's Temporary Chat and Microsoft Copilot's temporary chat show the same rule from another angle: some conversations should not become future personalization. Microsoft Copilot also provides deleting saved memory plus turning off personalization based on chat history.
For teams, the implementation lesson is sharper. Anthropic's memory tool documentation puts storage plus path restrictions on the application side, with just-in-time retrieval. In practice, ordinary workers should save memories that reduce repeated explanation. They should also save memories that block repeated errors or preserve agreed standards. Everything else should expire, stay project-bound, or require approval before it follows the user into the next task.
How to test memory before trusting it
A memory feature should earn trust by reducing re-explanation without turning old context into hidden policy. The first eval is simple: give the system a task after prior work and check whether it recalls the project's decisions, the reason a previous attempt failed, a customer's preferences, or acceptance criteria for that task type. Woshipm's author describes that as process learning. The harder part is checking whether the same memory stays quiet when it no longer applies.
That means testing both recall and restraint.
External memory is not one thing. Woshipm lists chat memory, project memory, RAG, knowledge bases, vector databases, and graph memory as separate forms. OpenAI's Memory FAQ also separates visible saved memories from a continuously updated synthesis drawn from past chat context. Those differences matter for eval design, because a user can review a listed memory more easily than a model's compressed sense of prior conversation.
A practical test suite should follow the lifecycle woshipm gives for continual learning. Capture asks whether the system noticed the right signal. Filtering asks whether it ignored noise. Consolidation asks whether it stored a usable summary rather than a transcript-shaped burden. Retrieval asks whether the memory arrived just in time. Forgetting asks whether correction actually changes later behavior; it also asks whether deletion changes later behavior. A reset must change later behavior too.
Project isolation needs its own test. Claude uses separate memory spaces and project summaries for each project, and separates project memory from non-project chats, according to woshipm. Anthropic's memory tool documentation also puts storage design on the application side: the app decides where data is stored, how it is stored, and how file paths are restricted. So the eval cannot stop at model answers. It must inspect boundaries.
Cost is broader than token spend. There is storage cost, retrieval cost, review cost, isolation cost, and reset cost. User controls are part of that accounting: OpenAI emphasizes reviewable memory summaries, Claude allows pause or reset and import or export across AI services, and Microsoft Copilot offers temporary chat, saved-memory deletion, and a personalization switch based on chat history. A memory that saves explanation but raises review burden may still be a bad trade.
Robots expose the same context problem in physical form
Robots make the context problem visible because the output is no longer a paragraph. It is a body moving through space. According to the report, Westlake Robotics began business in January 2024 and has already completed a Series A round, raising 500 million yuan in 6 months across 4 financing rounds. The money is aimed at research and development of a full-body unified large model for humanoid robots, plus an embodied intelligence talent hub.
Westlake's premise is control first. Wang Donglin argues that humanoid robots need general full-body humanoid motion control before higher-level perception is stacked on top. Its GAE model, described by the report as a Transformer-based general action large model, is meant to map user intent or instructions directly into stable full-body movement without presetting or programming. It is also pitched for robots from different manufacturers, including remote operation and dangerous-substitution scenarios.
That is one side of the bottleneck: turning intent into coordinated action.
Shutu Technology starts from the other side. Founded in 2024, it is described by the report as embodied intelligence infrastructure that connects professional human behavior with robot capabilities. Its pipeline converts human experience from real jobs into data and skills robots can learn, using physical information parsing, behavior representation, and human-robot action mapping. The company summarizes the route as "Human Action -> Learnable Physical Experience -> Robot Action."
The contradiction is useful for anyone designing AI workflows. Westlake is trying to solve the execution layer: how a machine should move when given an instruction. Shutu is trying to solve the memory layer: which traces of human work should become reusable machine context. Shutu also combines professional human behavior data with real robot data for training and validation, and the report says more than 60% of domestic embodied intelligence companies with valuations exceeding 10 billion yuan have become its customers.

Mu Wei's point sharpens the lesson. He sees data as only one visible contradiction, while the long-term problem is using data well, turning it into capability, and handling implementation operations and systemic problems. In workflow terms, context is not stored history. It is selected behavior, represented for action, retrieved when useful, and constrained by the task.
Context governance: what to capture, retrieve, and withhold
- You are building an agent that must improve across sessions or projects. Treat memory as a governed workflow, not a chat transcript. The source material describes usable continual learning as five stages: capture, filtering, consolidation, retrieval, and forgetting. Use external memory for items such as chat memory, project memory, RAG, knowledge bases, vector databases, and graph memory; reserve process learning for reusable decisions, failure reasons, customer preferences, or task acceptance criteria.
- You need persistent memory but want to avoid cross-project contamination. Use separated memory spaces or project-level summaries. Claude is described as having separate memory spaces and project summaries for each project, while OpenAI emphasizes viewable memory summaries that let users review, add, and update information. Claude also allows users to pause or reset memory, and supports importing and exporting memory across different AI services.
- A conversation should not shape future personalization. Offer an explicit non-memory mode. Gemini launched Temporary Chat alongside personal context, with the source noting that some chats should not enter future personalization and will not be used for training. Microsoft Copilot also provides controls including temporary chat, deleting saved memory, and turning off personalization based on chat history.
- You are adding multimodal input such as voice, gaze, gesture, touch, facial expression, spatial action, physiological state, or attention level. Add modalities only when they improve efficiency, accuracy, or fluency compared with a single modality. The source argues that multimodal input combinations should generally be restrained to 2-3 modalities unless there are special circumstances, and should adapt to physiology, cultural customs, and habit preferences.
- You are turning human behavior into training data for robots or embodied agents. Represent behavior as learnable physical experience, not just raw recordings. Shutu Technology's approach is summarized as "Human Action -> Learnable Physical Experience -> Robot Action," using physical information parsing, behavior representation, and human-robot action mapping, and combining professional human behavior data with real robot data for training and validation.
- You are considering implicit human-state signals such as physiology, attention, EEG, blood flow, ultrasound, MRI images, or behavioral data. Use them only when the task justifies the extra sensitivity and complexity. The sources distinguish explicit input from implicit input, and Gestala's work shows the frontier case: it is collaborating with embodied intelligence companies including Fourier to synchronously collect operators' EEG and cerebral blood-flow signals during embodied intelligence training, to test whether embodied models can gain better generalization by adding multimodal brain data than by using only action data.
Multimodal context needs consent, not demo logic
Multimodal context is attractive because it promises to turn weak signals into usable control. The woshipm author frames AI's shift from large language models to multimodal models and world models as a move from symbolic semantic parsing toward physical context understanding. Ben Shneiderman's definition makes the scope concrete: expression can include body movement, spoken language, or gaze. Perception can include hearing. It can also rely on vision or touch.
That expansion changes the consent problem.
A click or spoken command is explicit input. Woshipm separates that from implicit input: physiological state and attention level that a system perceives without the user actively expressing them. The difference matters for workflow design. A user can correct a command they gave. They may never know that a system interpreted their attention level, stored it, and used it to choose the next action. Context governance has to draw that boundary before capability demos blur it.
Brain-computer interface funding shows why the boundary is no longer theoretical. According to geekpark, Neuralink said in January 2026 that participants in its global clinical trials had increased to 21. OpenAI became the largest investor in Merge Labs' approximately $250 million seed round. Synchron completed a $200 million Series D round. Public statistics cited by geekpark put China's brain-computer interface sector at 17 financing events in the first quarter of 2026, totaling about 3.8 billion yuan.
Gestala makes the workflow question sharper. Geekpark reports that Peng Lei and Chen Tianqiao co-founded the company in January 2026, choosing ultrasound brain-computer interfaces, a field with few prior entrants. With embodied intelligence companies including Fourier, Gestala is collecting operators' EEG and cerebral blood-flow signals during training to test whether embodied models generalize better with multimodal brain data than with action data alone.
More signal is not automatically better context. Gestala's planned brain foundation model combines electrical signals, ultrasound, blood flow, MRI images and behavioral data. The woshipm author's rule is stricter: add multimodal input only when it improves efficiency or accuracy. Fluency can matter too, and combinations usually should stay at 2-3 modalities unless special circumstances justify more.
That is the consent standard agents need. Explicit input can drive action. Implicit input should require a higher bar: clear benefit plus limited capture. It also needs user adaptation and a way to correct or withdraw it. Without that, multimodal systems do not gain context governance. They gain surveillance with a demo interface.
Context governance patterns across agents, multimodal interfaces, robots and brain-computer systems
| Dimension | AI agents and memory products | Multimodal interface design | Embodied robot training | Brain-computer and brain-data systems |
|---|---|---|---|---|
| Context signals captured | External memory includes chat memory, project memory, RAG, knowledge bases, vector databases, and graph memory; process learning can record project decisions, failure reasons, customer preferences, or acceptance criteria. | Explicit input includes contact, touch, holding, voice, gaze, facial expressions, and spatial actions; implicit input includes physiological state and attention level. | Shutu Technology converts professional human behavior from real jobs into data and skills robots can learn; Westlake Robotics maps user intents or instructions into full-body movements. | Gestala plans to use electrical signals, ultrasound, blood flow, MRI images and behavioral data at the same time; it collaborates with embodied intelligence companies to collect operators' EEG and cerebral blood-flow signals. |
| How context is represented for machines | AI continual learning is described as parameter learning, external memory, and process learning; a usable continual learning system usually includes capture, filtering, consolidation, retrieval, and forgetting. | The woshipm author divides multimodal input into explicit input and implicit input and proposes a design framework of design motivation, design experience, and adaptation to change. | Shutu Technology uses physical information parsing, behavior representation, and human-robot action mapping, summarized as "Human Action -> Learnable Physical Experience -> Robot Action." GAE is based on a Transformer architecture. | Functional ultrasound imaging reads blood-flow changes to capture signals of neural activity, and transcranial focused ultrasound modulates neural activity in deep brain regions through phased arrays. |
| Retrieval or use at execution time | Anthropic's memory tool emphasizes just-in-time context retrieval; Claude reads, writes, and updates memory in real time during conversations. | The author argues multimodal input should only be introduced when it improves efficiency, accuracy, or fluency compared with a single modality. | Shutu Technology combines professional human behavior data with real robot data for training and validation; GAE can support remote operation and dangerous-substitution scenarios. | Gestala and partners are testing whether embodied models gain better generalization by adding multimodal brain data than by using only action data. |
| Human control, correction, forgetting or boundaries | OpenAI emphasizes viewable memory summaries; Claude allows users to pause or reset memory and separates project memory from non-project chats; Gemini launched Temporary Chat; Microsoft Copilot provides temporary chat, deleting saved memory, and turning off personalization based on chat history. | The author argues multimodal input combinations should generally be restrained to 2-3 modalities unless there are special circumstances, and should adapt to physiology, cultural customs, and habit preferences. | not covered | not covered |
| Stated bottleneck or design caution | The woshipm author argues AI is moving from generating content into systems that understand goals, connect resources, and execute tasks. | The author argues the essence of multimodal interaction is not adding modalities, but constructing a human-centered dialogue loop. | Mu Wei believes the long-term problem is how robots use data well and turn data into capabilities; Wang Donglin argues humanoid robots must first solve general full-body humanoid motion control. | Peng Lei said Gestala is decoding the brain while building a foundation model for the human brain. |
| Commercial or product status | Pricing not disclosed in sources. | Pricing not disclosed in sources. | Shutu Technology completed angel++ and angel+++ financing rounds; Westlake Robotics raised 500 million yuan in 6 months and completed 4 financing rounds within half a year. | Gestala completed a 420 million yuan angel+ financing round in July; its two financing rounds have raised nearly $100 million in total. |
Treat memory as a workflow surface you can inspect, not as a hidden model upgrade. Woshipm notes that OpenAI emphasizes viewable memory summaries. Claude can pause or reset memory, keep project memory separate from non-project chats, and import or export memory across AI services.
Start by marking which work should be remembered and which should stay temporary. Gemini's Temporary Chat and Microsoft Copilot's temporary chat point to a practical rule. So do saved-memory deletion and chat-history personalization controls. Sensitive drafts need an off-ramp, along with one-off negotiations and exploratory prompts.
For teams building agents, the open thread is storage discipline. Anthropic's documentation puts memory tools on the application side, with the application controlling where data sits, how it is stored, and how paths are restricted. Map your own capture and filtering before adding automation. Decide how consolidation should work. Specify retrieval rules, then define forgetting. Then test whether project context, past feedback, or a growing knowledge base actually improves the next task.
For readers outside China
- Availability: The source material is about Chinese and global product and research trends rather than a single purchasable tool. It discusses OpenAI, Anthropic, Google Gemini, Microsoft 365 Copilot, Shutu Technology, Westlake Robotics, and Gestala. Availability outside China is not disclosed in sources for Shutu Technology, Westlake Robotics, or Gestala.
- Pricing: Product pricing is not disclosed in sources. The sources do include financing figures: Shutu Technology's angel++ round raised tens of millions of yuan and its angel+++ round was nearly 100 million yuan; Westlake Robotics raised 500 million yuan in 6 months; Gestala completed a 420 million yuan angel+ financing round and its two financing rounds have raised nearly $100 million in total; Synchron completed a $200 million Series D financing round; OpenAI became the largest investor in Merge Labs' approximately $250 million seed financing round.
- Closest Western equivalents: ChatGPT memory, including saved memories and continuously updated synthesis from past chat context; Claude memory, project summaries, separated project memory, and memory import/export controls; Gemini personal context and Temporary Chat; Microsoft 365 Copilot Memory, temporary chat, saved-memory deletion, and personalization controls; RAG systems using knowledge bases, vector databases, or graph memory; Anthropic MCP, Computer Use, and Agent Skills for connecting models to tools and workflows
- Data residency: The sources do not provide data-residency guarantees or hosting locations for the memory and agent features discussed. Anthropic's memory tool documentation is described as application-side: the application controls where data is stored, how it is stored, and how paths are restricted. For Chinese embodied-intelligence companies, Gestala is described as having its global headquarters and manufacturing base in Chengdu and a second headquarters in Shanghai, but the source material does not specify where customer, robot, brain-signal, or behavioral datasets are stored.
Sources
- woshipm 从 ChatGPT 到 Agent:AI 每半年换一次牌桌,我们靠什么不掉队? https://woshipm.com/ai/6444079.html
- woshipm AI持续学习火了:下一代智能体拼的不是"会回答",而是"会积累" https://woshipm.com/ai/6438276.html
- 36kr 西湖大学教授创业做具身通用大脑,半年完成4轮共5亿元融资|硬氪首发 https://36kr.com/p/3924908647577732
- 36kr 硬氪首发 | 深圳具身基础设施公司获近亿元融资,拿下60%估值百亿头部机器人公司客户 https://36kr.com/p/3954612909227138
- geekpark 对话格式塔彭雷:AI 的下一站,是解读人的大脑 https://geekpark.net/news/368694
- woshipm 自然交流设计-多模态交互-上篇 https://woshipm.com/ucd/6439321.html
The evidence: 88 facts from 6 Chinese articles
Each line below was extracted from the article it sits under, in Chinese, before any of this was written. The writing is done from these and never from the source prose - that separation is structural, not a promise. How we work.
36kr硬氪首发 | 深圳具身基础设施公司获近亿元融资,拿下60%估值百亿头部机器人公司客户
- Shutu Technology completed angel++ and angel+++ financing rounds consecutively.
- Shutu Technology's angel++ financing round raised tens of millions of yuan.
- Shutu Technology's angel++ financing round included participation from Lingge Venture Capital, Shenzhen HTI Group, and an unnamed leading mobile phone and intelligent driving industry investor.
- Shutu Technology's angel+++ financing round was disclosed as nearly 100 million yuan.
- Shutu Technology's angel+++ financing round investors included Cornerstone Capital, Nanshan Strategic Emerging Industry Investment, Suzhou Trend Capital, Energy Conservation Capital, Hangshi Group, and West Lake Innovation Investment.
- Shutu Technology was founded in 2024.
- Shutu Technology's main business in the past was data collection.
- Shutu Technology wants to become an embodied intelligence infrastructure company.
- Shutu Technology uses physical information parsing, behavior representation, and human-robot action mapping to convert professional human behavior into physical experience that robots can learn.
- Shutu Technology combines professional human behavior data with real robot data for training and validation.
- Shutu Technology summarizes its approach as "Human Action -> Learnable Physical Experience -> Robot Action."
- Shutu Technology has formed a product matrix covering multiple source data forms.
36kr西湖大学教授创业做具身通用大脑,半年完成4轮共5亿元融资|硬氪首发
- Westlake Robotics recently completed a Series A financing round.
- Westlake Robotics raised 500 million yuan in 6 months.
- Westlake Robotics completed 4 financing rounds within half a year.
- Westlake Robotics' investors include SAIF, Xiaomiao Langcheng, Henan Investment Group Huirong Fund, and Haiyuan Fund.
- Westlake Robotics plans to focus the financing proceeds on research and development of a full-body unified large model for humanoid robots and on building an embodied intelligence talent hub.
- Westlake Robotics' business began in January 2024.
- Westlake Robotics is Westlake University's first technology transfer company in artificial intelligence and robotics.
- Westlake Robotics focuses on coordinated development of end-to-end embodied intelligence large models and robot bodies.
- Westlake Robotics launched its fully self-developed humanoid robot body named Westlake o1.
- Westlake Robotics founder Wang Donglin is a chief scientist for a National Science and Technology Innovation 2030 Major Project, a tenured professor at Westlake University, and deputy director of the Department of Artificial Intelligence at Westlake University.
- Wang Donglin is described as a 2026 AAAI Best Paper award winner.
- Westlake Robotics co-founder Zhang Yue is a tenured professor at Westlake University and vice dean of the School of Engineering at Westlake University.
- Westlake Robotics' core research and development members come from companies including Alibaba, ByteDance, Tencent, and Huawei, and universities including Oxford, Cambridge, Carnegie Mellon, Berkeley, Technical University of Munich, Tsinghua University, and Zhejiang University.
- Westlake Robotics' core research and development members have published more than 300 papers at top conferences including ICML, NeurIPS, CVPR, and RSS.
- GAE is a general action large model based on a Transformer architecture.
geekpark对话格式塔彭雷:AI 的下一站,是解读人的大脑
- In January 2026, Neuralink announced that the number of participants in its global clinical trials had increased to 21.
- In January 2026, OpenAI became the largest investor in Merge Labs' approximately $250 million seed financing round.
- Merge Labs was co-founded by Sam Altman and is betting on ultrasound rather than electrodes to read and modulate the brain without craniotomy.
- Synchron completed a $200 million Series D financing round for pivotal trials and commercialization preparations.
- Public statistics show that China's brain-computer interface sector had 17 financing events in the first quarter of 2026, with a total amount of about 3.8 billion yuan.
- Gestala completed a 420 million yuan angel+ financing round in July, after being founded only a little more than half a year earlier.
- Gestala's two financing rounds have raised nearly $100 million in total.
- Peng Lei and Chen Tianqiao co-founded Gestala in January 2026.
- Gestala chose ultrasound brain-computer interfaces, a field that previously had few entrants.
- Gestala built its global headquarters and manufacturing base in Chengdu and established a second headquarters in Shanghai for AI brain foundation model research and international scientific collaboration.
- Gestala is collaborating with embodied intelligence companies including Fourier to synchronously collect operators' EEG and cerebral blood-flow signals during embodied intelligence training.
- Peng Lei previously co-founded the invasive brain-computer interface company NeuroXess.
- Functional ultrasound imaging can read blood-flow changes to capture signals of neural activity, and transcranial focused ultrasound can modulate neural activity in deep brain regions through phased arrays.
- The name Gestala comes from the German philosophical and psychological concept Gestalt, whose core idea is that the whole is greater than the sum of its parts.
- Gestala's planned brain foundation model will use electrical signals, ultrasound, blood flow, MRI images and behavioral data at the same time.
- Clinical data in the United States showed that after ultrasound modulation of the anterior cingulate cortex, subjects' pain levels decreased by about half and the effect lasted one to two weeks.
- Gestala plans to release its first-generation model by the end of 2026, together with its first-generation ultrasound brain-computer interface product, preliminary clinical data on chronic pain and staged scientific research results.
woshipm自然交流设计-多模态交互·上篇
- Ben Shneiderman defines multimodal interaction as the process in which humans communicate with computer systems through multiple expression channels, such as body movement, spoken language, and gaze, and multiple perception channels, such as hearing, vision, and touch.
- Multimodal interaction was first introduced in 1979 in Richard A. Bolt's MIT work "Put-that-there: voice and gesture at the graphics interface."
- The woshipm author divides multimodal input into explicit input and implicit input based on the degree of user participation.
- Explicit input refers to conscious and active user interaction behaviors, such as contact, touch, holding, voice, gaze, facial expressions, and spatial actions.
- Implicit input refers to signals that users do not actively express and that systems automatically perceive, such as physiological state and attention level.
- The woshipm author proposes a multimodal input design framework of design motivation, design experience, and adaptation to change.
- Christopher Wickens's Multiple Resource Theory states that human cognitive and action resources are distributed across multiple relatively independent resource pools rather than one single whole.
- The woshipm author identifies physiology, cultural customs, and habit preferences as three dimensions that multimodal input design should adapt to for individual users.
woshipm从 ChatGPT 到 Agent:AI 每半年换一次牌桌,我们靠什么不掉队?
- The 1995 paper "Intelligent Agents: Theory and Practice" discussed the theory, architecture, and applications of intelligent agents.
- Lewis and others published a 2020 paper that systematically proposed retrieval-augmented generation.
- DALL-E 2 entered public testing in 2022.
- Early versions of Midjourney began to be used by many creators in 2022.
- Stable Diffusion was publicly released in 2022.
- OpenAI released the ChatGPT research preview on November 30, 2022.
- GPT-4 was released in March 2023.
- OpenAI added function calling capability to its API in June 2023.
- GPT-4o put text, images, and audio into more natural real-time interaction in 2024.
- Gemini 1.5 pushed long context into product discussions at the million-token level in 2024.
- Sora showed the possibility of generating long videos from text in 2024.
- Anthropic publicly tested Computer Use in October 2024.
- Anthropic open-sourced MCP in November 2024.
- DeepSeek-V3 released a technical report in December 2024.
- DeepSeek-R1 publicly released its model and paper in January 2025.
- OpenAI released the Responses API and Agents SDK in March 2025.
- OpenAI released the Codex research preview in May 2025.
- Anthropic launched Agent Skills in October 2025.
- OpenAI used Skills for ChatGPT, Codex, and API in 2026.
woshipmAI持续学习火了:下一代智能体拼的不是“会回答”,而是“会积累”
- Parameter learning includes continued pretraining, domain fine-tuning, and incremental training.
- External memory includes chat memory, project memory, RAG, knowledge bases, vector databases, and graph memory.
- OpenAI's ChatGPT memory updates described earlier saved memories as more like keeping a few notes.
- OpenAI introduced dreaming after 2025 to organize ChatGPT memory in the background by referring to chat history, rather than relying only on users explicitly saying "please remember."
- OpenAI's Memory FAQ says ChatGPT's memory is not only the saved memories shown in a list but also a continuously updated synthesis from past chat context.
- Claude's official help documentation says Claude organizes memory into categorized entries after memory is enabled and reads, writes, and updates memory in real time during conversations.
- Claude has separate memory spaces and project summaries for each project to avoid contamination across projects.
- Google Gemini launched personal context based on past chats in 2025.
- Google said Gemini can refer to past chats to learn preferences and provide more personalized responses.
- Microsoft 365 Copilot puts Memory into office scenarios so that Copilot can remember communication style, work goals, preferred topics, and recurring tasks.
- OpenAI's memory update emphasizes viewable memory summaries that let users review, add, and update information they want ChatGPT to know.
- Claude allows users to pause or reset memory, separates project memory from non-project chats, and supports importing and exporting memory across different AI services.
- Gemini launched Temporary Chat alongside personal context, indicating that some chats should not enter future personalization and will not be used for training.
- Microsoft Copilot provides controls including temporary chat, deleting saved memory, and turning off personalization based on chat history.
- Anthropic's memory tool documentation lets Claude save cross-session information through client-side file operations and emphasizes just-in-time context retrieval.
- Anthropic's memory tool documentation tells developers that memory tools are executed on the application side, and that the application controls where data is stored, how it is stored, and how paths are restricted.
- Claude memory tool use cases include maintaining project context across multiple Agent sessions, applying past interactions and feedback to new tasks, and building a knowledge base over time.