Paper Slam 4/19: When a Web Designer Meets a Radiologist — Two AI Agents, Two Paths
*English structured digest of a Chinese forum post comparing two agent papers (original by forum author “小凯”, written in a Feynman-style analytical voice).*
The two papers
- MM-WebAgent (arXiv 2604.15309, Microsoft Research Asia): a hierarchical multimodal web agent for webpage generation. Core challenge: creative integration — making AI-generated images, videos, and charts look coherent in one page.
- RadAgent (arXiv 2604.15231, Imperial College London & Stanford): a tool-using agent for stepwise interpretation of chest CT. Core challenge: explainable correctness — every conclusion must be traceable to imaging evidence.
- Problem: Directly collaging AIGC outputs (Midjourney images, Sora videos, charts) into a webpage produces visual inconsistency — each asset is fine alone, but together they clash.
- Architecture: A central LLM (gpt-4o) orchestrates AIGC tools via tool use, treating multimodal content generation as a first-class action alongside code writing.
- Three-stage planning: (1) *Global planning* — a structured JSON design blueprint (layout, style, color scheme, sections); (2) *Local planning* — per-section asset generation with style-matched prompts; (3) *Integration & reflection*.
- Three-level reflection: Local Refine (per-asset quality) → Context Refine (HTML/CSS layout fixes) → Global Refine (render, screenshot, whole-page aesthetic review).
- MM-WebGEN-Bench: a new benchmark of 120 design tasks spanning four intents, multiple visual styles, layout complexities, and multimodal combinations, scored on six dimensions: layout correctness, style coherence, aesthetics (global) and image/video/chart quality (local).
- Results (average score): Code-only One-shot 0.43 → Code-only Agent 0.49 → MM-WebAgent 0.75. Multimodal content generation roughly doubles local asset quality — the gap is categorical, not incremental. It is also competitive on WebGen-Bench (functional code).
- Limitations: quality depends on external AIGC tools; fixed tool set (no 3D generation, etc.); training-free with no learning from experience — a smart “process executor,” not an “experience accumulator.”
- Problem with end-to-end 3D VLMs (e.g., CT-CHAT): black-box reasoning, severe hallucination, adversarial fragility (macro-F1 only 0.287 on CT-RATE), and 0% faithfulness — report claims cannot be traced to image evidence.
- Approach: stepwise reasoning with interpretable tools — region localization, lesion detection, measurement, cross-slice comparison, report templating. Every tool call has human-readable inputs/outputs and is recorded in an inspectable trace.
- Faithfulness verification: each report assertion is matched against visual evidence. CT-CHAT: 0%; RadAgent: 37.0% — a from-zero-to-one breakthrough in verifiability.
- Results (CT-RATE, 50,188 chest CTs): macro-F1 0.287 → 0.347, micro-F1 0.312 → 0.366, precision and recall each up ~5.8 points. Adversarial robustness improves by 24.7 points, attributed to protocol-constrained tool calling acting as regularization. Recall gains matter most clinically (fewer missed findings).
- Limitations: 63% of report content still unverifiable; tool set is chest-CT specific; stepwise inference is far slower than end-to-end — problematic in time-critical settings like the ER.
- Naming ≠ understanding: both are “Agents,” but one is a creative coordinator and the other a reasoning executor. “Agent” is a spectrum, and the term is suffering inflation — many so-called agents are just chatbots with function calling. More precise descriptions carry more information.
- Demonstration over argument: MM-WebAgent persuades with side-by-side screenshots; RadAgent persuades with fully open reasoning traces.
- Assembly is the underrated skill: neither paper invents new AI capabilities; both demonstrate *organizational* ability — composing existing parts (LLM reasoning, AIGC generation, visual encoders) into structured workflows.
- Open questions: Is MM-WebAgent's lack of learning a flaw or a guard against repetitive “design ruts” (and where does machine “taste” come from — training data full of AI slop caps the ceiling)? Is 37% faithfulness enough? Probably not as a final diagnosis, but valuable as a triage tool that directs radiologist attention — AI as *augmentation*, not replacement.
- Honest caveats: hierarchical planning is standard software engineering; RadAgent's latency may rule it out in emergencies; neither paper addresses noisy/ambiguous inputs.
Key points — MM-WebAgent
Key points — RadAgent
Head-to-head insight
| Dimension | MM-WebAgent | RadAgent | |---|---|---| | Task | Webpage generation (creative) | CT report generation (medical) | | Core mechanism | Hierarchical planning + 3-level reflection | Stepwise tool use + faithfulness verification | | Key innovation | AIGC generation as a first-class action | Every reasoning step leaves a checkable trace | | Benchmark | MM-WebGEN-Bench (self-built) | CT-RATE (public) | | Cost of error | Aesthetic disaster (redoable) | Misdiagnosis (potentially fatal) |
Central thesis: the two agents occupy opposite poles of the agent spectrum, and the choice of philosophy is dictated by the cost of being wrong. Web design tolerates iterate-and-fix; medicine demands step-by-step verifiability. Neither approach is universally “better” — different domains require different constraints.
Feynman-style reflections from the author
Conclusion
Both systems share the same skeleton — a central LLM brain, specialized tools, and a feedback loop (visual for MM-WebAgent, logical for RadAgent). Together they sketch the future of agents: creative yet constrained, intelligent yet transparent, efficient yet trustworthy. For product builders: customer-service agents should follow RadAgent (every claim traceable); content-creation agents should follow MM-WebAgent (overall coherence beats individual brilliance).
*References: MM-WebAgent — arXiv 2604.15309 (Microsoft Research Asia); RadAgent — arXiv 2604.15231 (Imperial College London & Stanford).*