Easy AI Daily | 2025-11-26
Model Releases & Updates
Black Forest Labs Releases FLUX.2 Series
The FLUX.2 family includes four variants — Pro (API-only), Flex (quality/speed control), Dev (32B open weights), and Klein (open source, coming soon) — plus the FLUX.2 VAE. Dev weights are available on Hugging Face, supporting multi-reference image generation and 4K resolution output.- FLUX.2 official blog
- FLUX.2 Dev open weights
- Anthropic announcement
- Gemini 3 documentation
- scaling01 summary
- stuhlmueller research task evaluation
- EpochAIResearch benchmarks
- hendrycks comparison
- Replicate support
- Hugging Face pipeline
- Reddit discussion
- Reddit discussion
- r/ClaudeAI discussion
- r/GeminiAI chart controversy
- Perplexity announcement
- LMArena announcement
- OpenRouter announcement
- Unsloth announcement
- NVIDIA support docs
- Reddit discussion
- GPU MODE discussion
- iScienceLuvr overview
- iScienceLuvr overview
- Kangwook_Lee calibration
- MaartenSap CoT effects
- Psyche Office Hours: December 4 (Thursday) at 1PM EST on Discord — event link
- DSPy Pune Meetup: In-person gathering in Pune, India, announced via X — announcement
- MCP Dev Summit: Upcoming, though some members have scheduling conflicts — Discord discussion
Anthropic Releases Claude Opus 4.5
Claude Opus 4.5 delivers improved performance, excelling at coding and research tasks (paper QA, systematic reviews). Pricing drops to $5/million input tokens and $25/million output tokens, with tool calling and multi-turn conversation support.Google Gemini 3 Series Updates
The Gemini 3 API adds reasoning depth control, vision token budgets, and Thought Signatures. It achieved 93% on the GPQA Diamond benchmark and supports multimodal reasoning.AI Twitter Highlights
Claude Opus 4.5 Performance & Applications
Leads Gemini 3 Pro on the SWE-Bench Verified coding benchmark; 96.5% accuracy on research tasks (paper QA, systematic reviews); supports BrowseComp-Plus tool calling.Google Gemini 3 Benchmark Results
Gemini 3 Pro set a 93% record on GPQA Diamond, with standout organic chemistry performance. Compared to Claude Opus 4.5: similar text reasoning, stronger visual input, slightly weaker jailbreak robustness.FLUX.2 Ecosystem Integration
Day-one support from Replicate, Together AI, and Vercel AI Gateway; Hugging Face provides an open pipeline; OSTris AI offers day-0 inference/editing and LoRA training tools.AI Reddit Highlights
FP8 Reinforcement Learning on Consumer GPUs
Unsloth's FP8 RL achieves 1.4x training speedup and 60% memory savings, running Qwen3:4B on just 5GB VRAM.FLUX.2 Runs on 24GB VRAM
Users got FLUX.2 running on an RTX 4090 (24GB) using diffusers local deployment, 4-bit quantization, and remote text encoders.Non-Technical Subreddit Reactions to Claude Opus 4.5
Users report improved complex coding ability but a more cautious tone (e.g., "roughly correct"). Charts show it outperforming Opus 4.1, though the y-axis starting values sparked controversy.AI Discord Highlights
Claude Opus 4.5 Arrives on Perplexity Max
Perplexity Max subscribers can now use Claude Opus 4.5; performance details undisclosed, but users praise its coding and research capabilities.FLUX.2 Added to LMArena and OpenRouter
LMArena added flux-2-pro and flux-2-flex (text-to-image and editing; multi-turn generation disabled, editing added). OpenRouter lists FLUX.2 [pro] (frontier quality) and FLUX.2 [flex] (complex text/details).Unsloth FP8 RL with NVIDIA Support
Unsloth's FP8 RL claims 1.4x training speed and 60% memory savings; NVIDIA officially supports it on Blackwell RTX-50 and DGX Spark with setup docs.Hardware & Infrastructure
NVIDIA RTX GPU Pricing & Market Trends
Used RTX 3090 (24GB) sells around $750; RTX 4090 at $2,000–3,500, with users favoring the 3090's value. RTX PRO 6000 Blackwell dropped to $7,999.NVIDIA B200 Benchmark Leak
Leaked benchmarks show B200 under CUDA runtime and Torch 2.9.1+cu130 completing 16384x7168 matrix ops in 33.6±0.05 µs and 7168x4096 in 124±0.1 µs.Research & Development
NVIDIA Nemotron-Flash Optimizes Small Models
Using evolutionary search to find hybrid attention/operator combinations, Nemotron-Flash improves the accuracy-latency frontier: 5.5% higher accuracy than Qwen3-0.6B, 1.3–1.9x lower latency, and 45.6x higher throughput.DiP Pixel-Space Diffusion Enables Fast Inference
DiP uses a two-stage DiT backbone and Patch Detailer Head for 10x faster inference with only 0.3% parameter overhead, achieving FID 1.90 on ImageNet 256x256.LLM-as-a-Judge Calibration Research
Research shows most LLM-as-a-Judge results use biased estimators and evaluator error rates need calibration; CoT explanations may increase blind user trust and reduce error detection.Community & Events
*Source: Easy AI teaching project*