Easy AI Daily | November 27, 2025
A roundup of AI industry news, model releases, research, and community discussions.
Agents & Tool Ecosystem
Anthropic Releases Persistent Agent Patterns and MCP Tasks Update
Anthropic proposed persistent agent practices (state checkpoints, structured artifacts, etc.); MCP released SEP-1686 "tasks" supporting background long-running tasks; LangChain clarified its framework–runtime–harness stack, positioning LangGraph as a runtime.- Anthropic blog summary
- MCP tasks announcement
- LangChain stack clarification
- Technical deep dive
- Memory announcement
- Virtual try-on
- LisanBench results
- Code Arena leaderboard
- Context compression update
- ModelScope page
- Reddit discussion
- LMArena announcement
- Comparison images
- Technical overview
- Announcement
- Paper
- Summary
- Summary
- Demo
- PixelDiT: Dual Transformer (patch-level and pixel-level) achieving FID 1.61 on ImageNet 256x256 and GenEval 0.74. Paper
- Apple STARFlow-V: Normalizing-flow video generation supporting T2V/I2V/V2V with causal prediction and flow-score matching. Paper
- FLUX.2 Pro: Richer detail vs. FLUX 1 Pro, no more "plastic" look. Comparison
- Nano Banana 2: Improved structured image generation on StructBench; community prompt resources shared. Analysis | Resources
- Hugging Face downloads: Chinese models reached 17.1% of downloads, surpassing US models, led by DeepSeek and Qwen; multimodal models are popular. Overview
- METR: Considered the most trusted external evaluator by practitioners. Comment
- AI Security Institute: Published a positive-but-caveated case study evaluating whether Opus 4.5 would sabotage AI safety research. Thread
- Zhihu: Used Qwen2.5-VL-72B/3B with LoRA fine-tuning for multimodal recommendations, improving MMEB-eval-zh by 7.4%. Write-up
- New benchmarks: MultiPathQA (pathology navigation), MTBBench (oncology decisions), WER is Unaware (clinical ASR). Pathology | MTBBench | WER
- Z-Image-Turbo buzz: Users note near-Seedream-4.0 performance; 6B parameters suits local deployment. Reddit
- Opus 4.5 ports ZBar to Swift 6: Solved a long-standing bug where other models failed. Reddit
- SWE-bench chart debate: Opus 4.5 leads at 80.9%, but chart design criticized. Reddit
- AI progress charts: Thomas Pueyo's "fun toy to AGI" chart questioned for rigor. Reddit
- Memes: Ilya Sutskever scaling quotes, Grok 4.1 unhinged replies, Gemini 3 satire. Singularity
- LMArena: Flux 2 vs. Nano Banana Pro comparisons; SynthID prevents nerfing. Announcement
- Perplexity: Thiel/Palantir concerns, Nvidia–OpenAI bubble talk.
- Unsloth: ERNIE developer challenge, ES HyperScale CPU training, Qwen3 fine-tuning issues. Devpost
- Cursor: Haiku for docs vs. Composer-1 for code; linting red-squiggle issues.
- GPU MODE: Triton kernels, NVFP4_GEMV leaderboard, NVRAR for multi-node inference. Paper
- OpenAI: ChatGPT political bias debates, Nano Banana comics.
- LM Studio: API endpoint errors, image captioning fixes. Docs
- OpenRouter: Opus overload, DeepSeek R1 removal, model fallback bugs. Fallback docs
- Nous Research: Psyche office hours, Suno–Warner deal, Blackwell INT/FP performance.
- Eleuther: Multi-stage LLM hallucinations, SGD shuffling debates, Emergent Misalignment replication. Paper
- Latent Space: Claude Code Plan Mode upgrades, Jeff Dean's 15-year ML retrospective.
- Yannick Kilcher: Information retrieval lecture, curriculum learning debates. Lecture
- HuggingFace: Inference API concerns, RapidaAI open-source voice platform. Rapida
- Modular: MAX examples, Python-based MAX controversy, Mojo API regressions.
- tinygrad: TinyJit kernel replay, random function implementation. Tutorial
- Moonshot AI: Kimi limits, canvas over chatbots.
- DSPy: dspy-cli open-sourced with FastAPI and MCP support. Repo
- MCP Contributors: New protocol version, namespace collision discussions.
- Manus.im: API quota errors affecting ~500 users.
- aider: Benchmark updates, Opus 4.5 upgrade survey, Bedrock model errors.
Booking.com Deploys Production Agents for Customer Messaging
Booking.com built agents with LangGraph, Kubernetes, GPT-4 Mini, and Weaviate semantic search, processing tens of thousands of messages daily with a 70% satisfaction improvement.Perplexity Launches Memory and Virtual Try-On
Perplexity added user-level Memory (view/delete/disable) and a virtual try-on shopping feature.Model Updates & Performance
Claude Opus 4.5 Shines in Benchmarks
Opus 4.5 Thinking ranked first on LisanBench and topped Code Arena WebDev; the non-Thinking version showed regressions, with community reports of Python tool overuse. Claude.ai added automatic context compression.Alibaba Open-Sources Z-Image-Turbo Text-to-Image Model
A 6B-parameter model based on the Qwen3 4B text encoder, free on ModelScope and integrated into Hugging Face Diffusers, with performance approaching Seedream 4.0.FLUX.2 Series Released
FLUX.2 pro/flex models joined LMArena with improved visual quality, eliminating the "plasticky" look; competitive against Nano Banana Pro.EGGROLL Speeds Up Evolution Strategies
EGGROLL uses low-rank perturbations to accelerate evolution strategies, supporting 100k+ populations and stable pre-training of recurrent LMs in large discrete systems.dnet Tackles Apple Silicon Memory Limits
Dria's dnet uses distributed inference, disk streaming, and UMA scheduling to run over-memory models on Apple Silicon clusters, resolving OOM issues.Inference & Efficiency
LatentMAS Cuts Multi-Agent Communication Tokens
LatentMAS replaces text-based communication with latent vectors, reducing tokens by 70–84% and improving speed 4–4.3x without accuracy loss.Reasoning Trace Distillation Lowers Costs
Training a 12B model on gpt-oss traces reduces token usage 4x and avoids repeated inference.Multimodal & Generative Models
Open-Source Ecosystem & Evaluation
Reddit Highlights
Discord Community Discussions
*Source: Easy AI teaching project.*