English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily News Digest | November 26, 2025

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for November 26, 2025 covers major model releases and community updates in the AI industry. Black Forest Labs launched the FLUX.2 series (Pro, Flex, 32B open-weight Dev, and upcoming Klein) with multi-reference generation and 4K output. Anthropic released Claude Opus 4.5 with strong coding and research performance at lower pricing ($5/$25 per million tokens), while Google's Gemini 3 API added reasoning depth control and scored 93% on GPQA Diamond. Community highlights include Unsloth's FP8 reinforcement learning for consumer GPUs (1.4x speedup, 60% VRAM savings), FLUX.2 running on 24GB VRAM via 4-bit quantization, and Claude Opus 4.5 arriving on Perplexity Max, LMArena, and OpenRouter. Hardware news covers RTX GPU pricing trends and leaked NVIDIA B200 benchmarks. Research items include NVIDIA's Nemotron-Flash, the fast pixel-space diffusion model DiP, and LLM-as-a-Judge calibration findings.

Easy AI Daily News | November 26, 2025

A digest of AI industry developments for November 26, 2025, compiled by the Easy AI teaching project.

Model Releases and Updates

Black Forest Labs Releases FLUX.2 Series

The FLUX.2 family includes Pro (API-only), Flex (quality/speed control), Dev (32B open weights), and Klein (upcoming open-source release), plus the FLUX.2 VAE. Dev weights are live on Hugging Face, supporting multi-reference image generation and 4K resolution output.
  • FLUX.2 official blog
  • FLUX.2 Dev open weights
  • Anthropic Releases Claude Opus 4.5

    Claude Opus 4.5 shows improved performance in coding and research tasks (paper QA, systematic reviews). Pricing dropped to $5/million input tokens and $25/million output tokens, with tool calling and multi-turn conversation support.
  • Anthropic announcement
  • Google Gemini 3 Series Updates

    The Gemini 3 API adds reasoning depth control, visual token budgets, and Thought Signatures. Gemini 3 Pro achieved 93% on the GPQA Diamond benchmark with multimodal reasoning support.
  • Gemini 3 documentation
  • AI Twitter Highlights

  • Claude Opus 4.5 evaluation: Leads Gemini 3 Pro on SWE-Bench Verified; research task accuracy (paper QA, systematic reviews) reaches 96.5%; supports BrowseComp-Plus tool calling. (scaling01, stuhlmueller)
  • Gemini 3 benchmarks: New 93% record on GPQA Diamond, strong in organic chemistry. Compared with Claude Opus 4.5: comparable text reasoning, better visual input, slightly weaker jailbreak robustness. (EpochAIResearch, hendrycks)
  • FLUX.2 ecosystem: Day-one support from Replicate, Together AI, and Vercel AI Gateway; Hugging Face open-source pipeline; OSTris AI offering day-0 inference/editing and LoRA training tools. (replicate, huggingface)
  • AI Reddit Highlights

  • FP8 reinforcement learning on consumer GPUs: Unsloth's FP8 RL achieves 1.4x training speedup and 60% VRAM savings; Qwen3:4B runs on 5GB VRAM. (discussion)
  • FLUX.2 on 24GB VRAM: Users run FLUX.2 on an RTX 4090 using diffusers with 4-bit quantized models and remote text encoders. (discussion)
  • Community feedback on Claude Opus 4.5: Users report better handling of complex coding problems but more hedged response style (e.g., "roughly correct"). A benchmark chart showing Opus 4.5 beating Opus 4.1 sparked controversy over its y-axis starting value. (r/ClaudeAI, r/GeminiAI chart debate)
  • AI Discord Highlights

  • Claude Opus 4.5 on Perplexity Max: Available to Perplexity Max subscribers; users report strong coding and research performance. (Perplexity Max)
  • FLUX.2 on LMArena and OpenRouter: LMArena added flux-2-pro and flux-2-flex with text-to-image and editing; OpenRouter lists FLUX.2 [pro] (frontier quality) and FLUX.2 [flex] (complex text/details). (LMArena, OpenRouter)
  • Unsloth FP8 RL with NVIDIA support: NVIDIA officially supports FP8 RL on Blackwell RTX-50 and DGX Spark with setup documentation. (Unsloth, NVIDIA docs)
  • Hardware and Infrastructure

  • NVIDIA RTX GPU pricing trends: Used RTX 3090 (24GB) around $750; RTX 4090 at $2,000–3,500, with users viewing the 3090 as better value; RTX PRO 6000 Blackwell dropped to $7,999. (Reddit discussion)
  • Leaked NVIDIA B200 benchmarks: Under CUDA runtime and Torch 2.9.1+cu130, B200 completes a 16384x7168 matrix operation in 33.6±0.05 µs and a 7168x4096 operation in 124±0.1 µs. (GPU MODE discussion)
  • Research and Development

  • NVIDIA Nemotron-Flash: Uses evolutionary search to discover hybrid attention/operator combinations, improving the accuracy-latency frontier of small models — 5.5% higher accuracy than Qwen3-0.6B with 1.3–1.9x lower latency and 45.6x higher throughput. (overview)
  • DiP pixel-space diffusion model: Two-stage DiT backbone with a Patch Detailer Head enables 10x faster inference with 0.3% parameter overhead, reaching FID 1.90 on ImageNet 256x256. (overview)
  • LLM-as-a-Judge calibration: Research notes most LLM-as-a-Judge results use biased estimators and require calibrating evaluator error rates; CoT explanations may increase blind trust and reduce error detection. (calibration method, CoT effects)
  • Community and Events

  • Psyche Office Hours: December 4 (Thursday) at 1PM EST on Discord. (event)
  • DSPy Pune meetup: In-person gathering in Pune, India; details announced on X. (announcement)
  • MCP Dev Summit: Upcoming, though some members cannot attend due to schedule conflicts. (discussion)
---

*Source: Easy AI teaching project.*

Tags

#ai-news#flux-2#claude-opus-4-5#gemini-3#unsloth#nvidia#diffusion-models#daily-digest

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169158