Easy AI Daily News Digest — November 26, 2025
*Translated and edited from the zhichai.net forum post (source: Easy AI teaching project).*
Model Releases and Updates
Black Forest Labs Releases FLUX.2 Series
The FLUX.2 family ships in four tiers: Pro (API-only), Flex (quality/speed control), Dev (32B open weights), and Klein (open weights coming soon), plus a FLUX.2 VAE. The Dev weights are live on Hugging Face, supporting multi-reference-image generation and 4K resolution output.- FLUX.2 official blog
- FLUX.2 Dev open weights
- Anthropic announcement
- Gemini 3 documentation
- Claude Opus 4.5 evaluation: Leads Gemini 3 Pro on SWE-Bench Verified; 96.5% accuracy on research tasks (paper QA, systematic reviews); supports BrowseComp-Plus tool calling. (scaling01, stuhlmueller)
- Gemini 3 benchmarks: 93% record on GPQA Diamond, strong in organic chemistry; comparable text reasoning to Opus 4.5, better visual input, slightly weaker jailbreak robustness. (EpochAIResearch, hendrycks)
- FLUX.2 ecosystem: Day-one support on Replicate, Together AI, Vercel AI Gateway; open-source pipeline on Hugging Face; OSTris AI offers day-0 inference/editing and LoRA training tools. (replicate, huggingface)
- FP8 reinforcement learning on consumer GPUs: Unsloth reports 1.4x training speedup and 60% VRAM savings; Qwen3:4B runs on 5GB VRAM. (discussion)
- FLUX.2 on 24GB VRAM: Community members run FLUX.2 on an RTX 4090 using diffusers and 4-bit quantization with remote text encoders. (discussion)
- Non-technical subreddit feedback on Opus 4.5: Users note improved complex coding ability but a more hedged response style; a benchmark chart comparing it favorably to Opus 4.1 drew criticism over its y-axis truncation. (ClaudeAI, GeminiAI)
- Claude Opus 4.5 on Perplexity Max: Available to Max subscribers; detailed specs undisclosed, but users report strong coding/research performance. (announcement)
- FLUX.2 on LMArena and OpenRouter: LMArena added flux-2-pro and flux-2-flex (text-to-image and editing); OpenRouter lists FLUX.2 [pro] and FLUX.2 [flex]. (LMArena, OpenRouter)
- Unsloth FP8 RL with NVIDIA support: NVIDIA officially supports FP8 RL on Blackwell RTX-50 and DGX Spark with setup docs. (Unsloth, docs)
- NVIDIA RTX GPU pricing trends: Used RTX 3090 (24GB) around $750; RTX 4090 at $2,000–3,500, with users favoring the 3090's value; RTX PRO 6000 Blackwell dropped to $7,999. (discussion)
- Leaked NVIDIA B200 benchmarks: Under CUDA runtime with Torch 2.9.1+cu130, a 16384x7168 matrix multiply took 33.6±0.05 µs; 7168x4096 took 124±0.1 µs. (GPU MODE thread)
- NVIDIA Nemotron-Flash: Uses evolutionary search over hybrid attention/operator mixes to improve the small-model accuracy–latency frontier: +5.5% accuracy over Qwen3-0.6B with 1.3–1.9x lower latency and 45.6x higher throughput. (overview)
- DiP pixel-space diffusion: Two-stage DiT backbone with a Patch Detailer Head achieves 10x faster inference with only 0.3% parameter overhead; FID 1.90 on ImageNet 256x256. (overview)
- LLM-as-a-Judge calibration: Most LLM-as-a-Judge results use biased estimators and require calibrating evaluator error rates; CoT explanations may increase blind trust and reduce error detection. (Kangwook Lee, Maarten Sap)
- Psyche Office Hours: December 4 (Thursday), 1PM EST on Discord. (event)
- DSPy Pune meetup: In-person gathering in Pune, India, announced via X. (announcement)
- MCP Dev Summit: Upcoming, though some community members cannot attend due to scheduling conflicts. (discussion)