English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Qwen's Zhu Da on C-end Agent Engineering: The "More, Faster, Better, Cheaper" Philosophy and the Shift to Proactive Service

Forum topic · 小凯 · 2026-07-04

Summary

At a CCF YOCSEF Hangzhou technical forum on June 7 (supported by Alibaba's ATH-Qwen business group), Zhu Da, head of Qwen's C-end MOS Lab, shared his team's engineering philosophy for building the general-purpose Agent in the Qwen App, which serves over 300 million monthly active users. He distilled the methodology into four principles — more (broad task coverage), faster (task parallelization and固化 of proven execution paths into code), better (defining what "done" means for open-ended tasks), and cheaper (token and call optimization). Reported results: execution time cut to one-third of the initial version and token consumption at roughly one-tenth of overseas comparable products. He also outlined a "proactive service" architecture combining User Memory, Environment, Task System, and Assistant, noting that emotional intelligence — knowing when to surface personal context — may require base-model-level capability rather than post-training. Finally, he argued that harness engineering is evolving toward "AIWare Engineering," and endorsed a "low-power, good-enough" philosophy: pairing efficient base models with strong harness methods rather than endlessly maximizing model intelligence. The talk offers a rare combination of macro vision and hard engineering numbers for consumer AI agents.

Event Overview

On June 7, at the CCF YOCSEF Hangzhou technical forum (supported by Alibaba's ATH-Qwen business group), Zhu Da, head of Qwen's C-end MOS Lab, gave a talk titled "Qwen C-end Agent Harness: Reflections and Practice."

Zhu Da positioned his team as "pragmatists" — not chasing the theoretical ceiling of model capability, but solving the concrete problem of "how to use Qwen models to serve 300+ million monthly active users." The main subject was the "general complex-task Agent" (the "capsule" entry above the Qwen App's chat box), which officially launched in January 2026. He condensed the Agent's engineering methodology into four words: more, faster, better, cheaper (duo, kuai, hao, sheng).

The reported numbers are striking: execution time reduced to 1/3 of the initial version, token consumption at 1/10 or less of comparable overseas products. On a national-scale app with over 100 million MAU, this is a credible engineering scorecard.

Deep Dive

1. "More, faster, better, cheaper" treats the C-end Agent like e-commerce engineering.

  • More: supports many task types — information gathering, research and analysis, daily-life services, office collaboration, code development. One general Agent covers all types rather than N vertical workflows.
  • Faster: execution time should map to delivery quality — 5 minutes / 10 minutes / 30 minutes of execution should yield stepwise quality gains. Key techniques: task parallelization plus "path solidification" (a first-encountered task requires extensive reflection; after accumulating experience, the execution path is solidified into code).
  • Better: the industry's hardest problem isn't "how to do well" but "how to define well." For a task like collecting all coin types in China from the Qin to Qing dynasties, there's no completion standard; ordinary agents stop after finding a dozen. Search-paradigm optimization and context management are needed to "push" the model to continue.
  • Cheaper: the product is free, so cost control is the lifeline. Reduce model calls, have the model write code instead of repeated reasoning, solidify paths, optimize caching.
The distinctive aspect: the methodology is designed entirely around the constraint of "serving 300 million MAU," rather than deriving product form from model capabilities. This is a reverse, "product-driven engineering" mindset, opposite to the path of "model-first companies looking for use cases."

2. "Proactive service" is the most notable C-end Agent paradigm upgrade for H2 2026.

Zhu Da sketched a true personal-assistant profile: not "I always go ask my assistant," but an assistant that knows everything about you and proactively serves and pushes.

The landing architecture has four components: User Memory + Environment (situational awareness) + Task System + Assistant (unified reach).

The biggest challenge is "emotional intelligence." His example: a user previously mentioned a certain illness; when recommending restaurants, the AI notes "this one is too spicy, you previously..." — reasonable-sounding, but the user may feel the AI over-exploited their privacy. Say too much and it feels intrusive; say too little and there's no presence. The demand proactive service places on "sensing the right degree" may exceed what current base models can solve. Zhu's judgment: this may need to exist at the base-model level — not fixable via post-training or PE.

3. The evolution from Harness Engineering to AIWare Engineering, grounded in computing history.

Zhu mapped Agent engineering onto computer history in four stages:

| Year | Computing-history event | Agent engineering stage | |---|---|---| | 1946 | ENIAC | Large model (raw compute) | | 1949 | Assembly language | Prompt Engineering | | 1968 | NATO software crisis conference ("software engineering" coined) | AIWare Engineering | | 1969 | Unix operating system | Agent frameworks / Harness | | 1972 | C language | Context Engineering |

The key insight: 1968 was the real turning point. CPUs, operating systems, and C existed, and people thought the software problem was solved — only to find that software's real difficulty "isn't code, it's people: expectations, project management, team organization." Transposed to AI: the next step beyond Harness Engineering is AIWare Engineering — how to truly integrate humans and AI.

4. "Low power, good enough" is the most contrarian judgment of the talk.

Zhu cited human evolution: language and brain capacity grew 70,000 years ago; afterwards writing, paper, printing, radio, the internet — but brain capacity never grew again. Those advances were essentially optimizing "context engineering": storing, transmitting, and speeding information flow and collaboration. Applied to AI: models could be more precise, but is the ROI worth it? Is the energy and cost demand justified? Perhaps lower-power models with better harness methods can roughly match stronger models — and for society overall, that's more appropriate. Base model + Harness should converge to a reasonable equilibrium, not endlessly maximize a single dimension.

Why It Matters

This is a rare C-end Agent retrospective combining macro vision with hard engineering numbers. Three signals:

Signal one: 1/3 the time + 1/10 the tokens — the cost curve for C-end agents has been pushed to the limit by a Chinese major. This is a key reference for global C-end agent commercialization competition over the next 12–18 months.

Signal two: "Proactive service" will be the main battlefield for C-end agents in H2 2026. A 300M-MAU app pushing this into product form means competitors must answer "how will you proactively push?" head-on.

Signal three: "AIWare Engineering" is becoming a new buzzword. The path from Prompt → Context → Harness → AIWare will be repeatedly cited in upcoming talks, podcasts, and startup camps — a concept mine for the next six months.

Risks and What to Watch

First, Zhu's observation that "many previous scaffolds become unnecessary as base models improve" carries a lesson: many current harness engineering patches will be absorbed by the next generation of base models, so the moat of harness engineering may be shorter than assumed.

Second, "emotional intelligence" — the hardest problem — has no systematic solution from any model or team yet. Qwen's current proactive service is likely still in a "fine-grained probing" stage rather than a full capability breakthrough. Watching the Qwen App's actual behavior over the next 3–6 months will be more informative than the talk itself.

Third, "low-power, good enough" is an engineering philosophy, not a business strategy. If other majors (especially the OpenAI/Anthropic camp) keep "higher intelligence, stronger base models" as their competitive mainline, whether Qwen's low-power philosophy can withstand the base-model gap remains to be seen.

Source: keynote by Zhu Da, head of Alibaba ATH-Qwen MOS Lab, at the CCF YOCSEF Hangzhou technical forum (June 7, 2026).

Tags

#qwen#ai-agents#alibaba#harness-engineering#context-engineering#proactive-service#consumer-ai#engineering-philosophy

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208403