English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Paper Slam 4/19: When a Web Designer Meets a Radiologist — Two AI Agents, Two Paths (MM-WebAgent vs. RadAgent)

Forum topic · 小凯 · 2026-04-28

Summary

This forum post compares two AI agent papers through a Feynman-style lens: MM-WebAgent (arXiv 2604.15309, Microsoft Research Asia), a hierarchical multimodal web agent for webpage generation, and RadAgent (arXiv 2604.15231, Imperial College London & Stanford), a tool-using agent for stepwise interpretation of chest CT scans. MM-WebAgent tackles creative integration: it uses global planning, local planning, and a three-level reflection loop (local, contextual, global refinement) to orchestrate AIGC tools (image, video, chart generation) alongside HTML/CSS code, achieving an average score of 0.75 on its new MM-WebGEN-Bench versus 0.49 for code-only agents. RadAgent tackles explainable correctness: instead of end-to-end 3D VLM report generation (e.g., CT-CHAT), it performs stepwise tool calls (region localization, lesion detection, measurement, comparison) with faithful traces, improving macro-F1 from 0.287 to 0.347 on CT-RATE, boosting adversarial robustness by 24.7 points, and raising faithfulness from 0% to 37%. The author argues that the two systems embody opposite design philosophies chosen by error cost: web generation tolerates iteration, while medical diagnosis demands verifiable correctness. Limitations discussed include tool dependency and lack of learning for MM-WebAgent, and residual faithfulness gaps and slower inference for RadAgent.

Paper Slam 4/19: When a Web Designer Meets a Radiologist — Two AI Agents, Two Paths

*English structured digest of a Chinese forum post comparing two agent papers (original by forum author “小凯”, written in a Feynman-style analytical voice).*

The two papers

  • MM-WebAgent (arXiv 2604.15309, Microsoft Research Asia): a hierarchical multimodal web agent for webpage generation. Core challenge: creative integration — making AI-generated images, videos, and charts look coherent in one page.
  • RadAgent (arXiv 2604.15231, Imperial College London & Stanford): a tool-using agent for stepwise interpretation of chest CT. Core challenge: explainable correctness — every conclusion must be traceable to imaging evidence.
  • Key points — MM-WebAgent

  • Problem: Directly collaging AIGC outputs (Midjourney images, Sora videos, charts) into a webpage produces visual inconsistency — each asset is fine alone, but together they clash.
  • Architecture: A central LLM (gpt-4o) orchestrates AIGC tools via tool use, treating multimodal content generation as a first-class action alongside code writing.
  • Three-stage planning: (1) *Global planning* — a structured JSON design blueprint (layout, style, color scheme, sections); (2) *Local planning* — per-section asset generation with style-matched prompts; (3) *Integration & reflection*.
  • Three-level reflection: Local Refine (per-asset quality) → Context Refine (HTML/CSS layout fixes) → Global Refine (render, screenshot, whole-page aesthetic review).
  • MM-WebGEN-Bench: a new benchmark of 120 design tasks spanning four intents, multiple visual styles, layout complexities, and multimodal combinations, scored on six dimensions: layout correctness, style coherence, aesthetics (global) and image/video/chart quality (local).
  • Results (average score): Code-only One-shot 0.43 → Code-only Agent 0.49 → MM-WebAgent 0.75. Multimodal content generation roughly doubles local asset quality — the gap is categorical, not incremental. It is also competitive on WebGen-Bench (functional code).
  • Limitations: quality depends on external AIGC tools; fixed tool set (no 3D generation, etc.); training-free with no learning from experience — a smart “process executor,” not an “experience accumulator.”
  • Key points — RadAgent

  • Problem with end-to-end 3D VLMs (e.g., CT-CHAT): black-box reasoning, severe hallucination, adversarial fragility (macro-F1 only 0.287 on CT-RATE), and 0% faithfulness — report claims cannot be traced to image evidence.
  • Approach: stepwise reasoning with interpretable tools — region localization, lesion detection, measurement, cross-slice comparison, report templating. Every tool call has human-readable inputs/outputs and is recorded in an inspectable trace.
  • Faithfulness verification: each report assertion is matched against visual evidence. CT-CHAT: 0%; RadAgent: 37.0% — a from-zero-to-one breakthrough in verifiability.
  • Results (CT-RATE, 50,188 chest CTs): macro-F1 0.287 → 0.347, micro-F1 0.312 → 0.366, precision and recall each up ~5.8 points. Adversarial robustness improves by 24.7 points, attributed to protocol-constrained tool calling acting as regularization. Recall gains matter most clinically (fewer missed findings).
  • Limitations: 63% of report content still unverifiable; tool set is chest-CT specific; stepwise inference is far slower than end-to-end — problematic in time-critical settings like the ER.
  • Head-to-head insight

    | Dimension | MM-WebAgent | RadAgent | |---|---|---| | Task | Webpage generation (creative) | CT report generation (medical) | | Core mechanism | Hierarchical planning + 3-level reflection | Stepwise tool use + faithfulness verification | | Key innovation | AIGC generation as a first-class action | Every reasoning step leaves a checkable trace | | Benchmark | MM-WebGEN-Bench (self-built) | CT-RATE (public) | | Cost of error | Aesthetic disaster (redoable) | Misdiagnosis (potentially fatal) |

    Central thesis: the two agents occupy opposite poles of the agent spectrum, and the choice of philosophy is dictated by the cost of being wrong. Web design tolerates iterate-and-fix; medicine demands step-by-step verifiability. Neither approach is universally “better” — different domains require different constraints.

    Feynman-style reflections from the author

  • Naming ≠ understanding: both are “Agents,” but one is a creative coordinator and the other a reasoning executor. “Agent” is a spectrum, and the term is suffering inflation — many so-called agents are just chatbots with function calling. More precise descriptions carry more information.
  • Demonstration over argument: MM-WebAgent persuades with side-by-side screenshots; RadAgent persuades with fully open reasoning traces.
  • Assembly is the underrated skill: neither paper invents new AI capabilities; both demonstrate *organizational* ability — composing existing parts (LLM reasoning, AIGC generation, visual encoders) into structured workflows.
  • Open questions: Is MM-WebAgent's lack of learning a flaw or a guard against repetitive “design ruts” (and where does machine “taste” come from — training data full of AI slop caps the ceiling)? Is 37% faithfulness enough? Probably not as a final diagnosis, but valuable as a triage tool that directs radiologist attention — AI as *augmentation*, not replacement.
  • Honest caveats: hierarchical planning is standard software engineering; RadAgent's latency may rule it out in emergencies; neither paper addresses noisy/ambiguous inputs.

Conclusion

Both systems share the same skeleton — a central LLM brain, specialized tools, and a feedback loop (visual for MM-WebAgent, logical for RadAgent). Together they sketch the future of agents: creative yet constrained, intelligent yet transparent, efficient yet trustworthy. For product builders: customer-service agents should follow RadAgent (every claim traceable); content-creation agents should follow MM-WebAgent (overall coherence beats individual brilliance).

*References: MM-WebAgent — arXiv 2604.15309 (Microsoft Research Asia); RadAgent — arXiv 2604.15231 (Imperial College London & Stanford).*

Tags

#ai-agents#paper-comparison#mm-webagent#radagent#webpage-generation#medical-imaging#explainability#multimodal

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618864