Easy AI Daily News | November 26, 2025
A digest of AI industry developments for November 26, 2025, compiled by the Easy AI teaching project.
Model Releases and Updates
Black Forest Labs Releases FLUX.2 Series
The FLUX.2 family includes Pro (API-only), Flex (quality/speed control), Dev (32B open weights), and Klein (upcoming open-source release), plus the FLUX.2 VAE. Dev weights are live on Hugging Face, supporting multi-reference image generation and 4K resolution output.- FLUX.2 official blog
- FLUX.2 Dev open weights
- Anthropic announcement
- Gemini 3 documentation
- Claude Opus 4.5 evaluation: Leads Gemini 3 Pro on SWE-Bench Verified; research task accuracy (paper QA, systematic reviews) reaches 96.5%; supports BrowseComp-Plus tool calling. (scaling01, stuhlmueller)
- Gemini 3 benchmarks: New 93% record on GPQA Diamond, strong in organic chemistry. Compared with Claude Opus 4.5: comparable text reasoning, better visual input, slightly weaker jailbreak robustness. (EpochAIResearch, hendrycks)
- FLUX.2 ecosystem: Day-one support from Replicate, Together AI, and Vercel AI Gateway; Hugging Face open-source pipeline; OSTris AI offering day-0 inference/editing and LoRA training tools. (replicate, huggingface)
- FP8 reinforcement learning on consumer GPUs: Unsloth's FP8 RL achieves 1.4x training speedup and 60% VRAM savings; Qwen3:4B runs on 5GB VRAM. (discussion)
- FLUX.2 on 24GB VRAM: Users run FLUX.2 on an RTX 4090 using diffusers with 4-bit quantized models and remote text encoders. (discussion)
- Community feedback on Claude Opus 4.5: Users report better handling of complex coding problems but more hedged response style (e.g., "roughly correct"). A benchmark chart showing Opus 4.5 beating Opus 4.1 sparked controversy over its y-axis starting value. (r/ClaudeAI, r/GeminiAI chart debate)
- Claude Opus 4.5 on Perplexity Max: Available to Perplexity Max subscribers; users report strong coding and research performance. (Perplexity Max)
- FLUX.2 on LMArena and OpenRouter: LMArena added flux-2-pro and flux-2-flex with text-to-image and editing; OpenRouter lists FLUX.2 [pro] (frontier quality) and FLUX.2 [flex] (complex text/details). (LMArena, OpenRouter)
- Unsloth FP8 RL with NVIDIA support: NVIDIA officially supports FP8 RL on Blackwell RTX-50 and DGX Spark with setup documentation. (Unsloth, NVIDIA docs)
- NVIDIA RTX GPU pricing trends: Used RTX 3090 (24GB) around $750; RTX 4090 at $2,000–3,500, with users viewing the 3090 as better value; RTX PRO 6000 Blackwell dropped to $7,999. (Reddit discussion)
- Leaked NVIDIA B200 benchmarks: Under CUDA runtime and Torch 2.9.1+cu130, B200 completes a 16384x7168 matrix operation in 33.6±0.05 µs and a 7168x4096 operation in 124±0.1 µs. (GPU MODE discussion)
- NVIDIA Nemotron-Flash: Uses evolutionary search to discover hybrid attention/operator combinations, improving the accuracy-latency frontier of small models — 5.5% higher accuracy than Qwen3-0.6B with 1.3–1.9x lower latency and 45.6x higher throughput. (overview)
- DiP pixel-space diffusion model: Two-stage DiT backbone with a Patch Detailer Head enables 10x faster inference with 0.3% parameter overhead, reaching FID 1.90 on ImageNet 256x256. (overview)
- LLM-as-a-Judge calibration: Research notes most LLM-as-a-Judge results use biased estimators and require calibrating evaluator error rates; CoT explanations may increase blind trust and reduce error detection. (calibration method, CoT effects)
- Psyche Office Hours: December 4 (Thursday) at 1PM EST on Discord. (event)
- DSPy Pune meetup: In-person gathering in Pune, India; details announced on X. (announcement)
- MCP Dev Summit: Upcoming, though some members cannot attend due to schedule conflicts. (discussion)
Anthropic Releases Claude Opus 4.5
Claude Opus 4.5 shows improved performance in coding and research tasks (paper QA, systematic reviews). Pricing dropped to $5/million input tokens and $25/million output tokens, with tool calling and multi-turn conversation support.Google Gemini 3 Series Updates
The Gemini 3 API adds reasoning depth control, visual token budgets, and Thought Signatures. Gemini 3 Pro achieved 93% on the GPQA Diamond benchmark with multimodal reasoning support.AI Twitter Highlights
AI Reddit Highlights
AI Discord Highlights
Hardware and Infrastructure
Research and Development
Community and Events
*Source: Easy AI teaching project.*