Easy AI Daily News | 2025-11-20
Key Points
Model Updates and Releases
- Google Gemini 3 Pro Image (Nano Banana Pro): Supports Google Search grounding, 2-4K resolution, and text-in-image generation/editing. Pricing: $0.134 per 2K image, $0.24 per 4K image. Available on the Gemini App/API, LM Arena, Hugging Face Spaces, and Together AI. Early demos show accurate infographics and chart annotation; text rendering error rate dropped from 56% to 8%. Includes SynthID watermarking.
- Links: Pricing details | Announcement | LM Arena | Hugging Face Spaces | Together AI | Error rate data | SynthID watermark
- AI2 Olmo 3: Fully open-source (Apache-2.0) with a 32B Think variant for long chain-of-thought and complex reasoning. Retains post-norm architecture; 7B uses sliding-window attention for KV cache optimization, 32B uses GQA. RL infrastructure delivers a 4x experiment speedup; emphasis on decontaminated evaluation (e.g., random reward tests).
- Links: Announcement reaction | Architecture analysis | RL infrastructure
- Meta SAM3 and SAM3D: SAM3 unifies image/video segmentation with text/visual prompting, 2x performance improvement, 30ms inference. SAM3D enables 3D reconstruction from a single image. Data engine contains 4M phrases and 52M masks; source code allows commercial use.
- Links: SAM3 announcement | SAM3D announcement | GitHub
- OpenAI GPT-5.1 Codex Max: Designed for long-running, detail-heavy tasks; first native multi-context-window support via compaction. SOTA on SWE-Bench; available only through ChatGPT plans, no API.
- Links: Release blog | Twitter announcement
- Cogito 2.1 in WebDev Arena: Deep Cogito's model ranks 18th overall, top-10 among open models. Hosted on Together and Fireworks; improvement details undisclosed.
- Links: Model page | WebDev Leaderboard
- OpenAI GPT-5.1 for Science: 13 early experiments shared, showing GPT-5.1 accelerating research in math, physics, biology, and materials science; 4 experiments helped solve previously unsolved problems.
- Links: Overview | arXiv paper thread
- Perplexity Comet browser: Released for Android, Mac, and Windows with voice-first browsing; supports Kimi-K2 Thinking and Gemini 3 Pro. Pro/Max users can create slides, tables, and documents.
- Cursor beta debugging mode: New log-ingest server automatically instruments code to collect logs; the agent validates hypotheses from logs rather than guessing.
- MemMachine Playground: Open Hugging Face Space supporting GPT-5, Claude 4.5, and Gemini 3 Pro with persistent AI memory; fully open-source.
- DSPy Proxy: New repo aryaminus/dspy-proxy simplifies DSPy agent development, built via a single prompt with Gemini 3 Pro.
- Jetson Spark cluster: A user built a 6-device NVIDIA Jetson cluster for NCCL/NVIDIA development, prototyping pre-B300 cluster workflows.
- Link: Reddit discussion
- GPU MODE discussions: GEMM optimization, CUDA caches (texture vs constant), AMD MI300X DMA collectives (+16% for large transfers), BF16 conversion issues (missing TensorRT kernels).
- Links: GEMM optimization article | DMA paper
- Mojo 0.25.7 regression: Nightly build throughput on Mac M1 running llama2.mojo dropped from ~1000 tok/sec to ~170 tok/sec; compiler team asked to investigate.
- Community explored jailbreaks for Gemini 3 Pro, Grok (shell access obtained), and Claude 4.5 (trust-building bypass producing methamphetamine synthesis steps).
- SynthID bypass: Users found a "do nothing" prompt via reve-edit can strip Gemini's SynthID watermark, or the model can be asked directly whether content is AI-generated.
- LMArena: Debates on Nano Banana Pro text rendering, SynthID bypasses, GPT-5.1 vs Gemini 3 Pro, Cogito 2.1's WebDev Arena performance.
- Perplexity community: Gemini 3 Pro coding praised over Claude Sonnet 4.5; Comet RAM usage concerns; Antigravity app called a "Cursor Killer."
- LM Studio: EmbeddingGemma recommended for RAG, Qwen3 thinking control, Mi60 GPU value, Vulkan crashes from model offloading.
- Unsloth AI: Gemini 3 Chrome integration speed vs local models; Cogito GGUF downloads; RAM price surge (64GB at $400).
- Yannick Kilcher: Skyfall AI's AI CEO benchmark (LLMs lag humans in long-horizon planning), SAM3D vs DeepSeek, NVIDIA Q3 earnings.
- Moonshot AI Kimi K2: Coding plan pricing ($19) seen as too high; SGLang tool-calling issues.
- Eleuther AI: KNN vs quadratic attention, Seth's conjecture, softmax attention distributions, IntologyAI's RE-Bench results claiming superhuman-expert performance.
- Nous Research: Gemma 3 hype (not AGI), World Models releases planned by DeepSeek/Qwen/Kimi.
- SAM3's unified architecture improves segmentation interpretability with text/visual prompting.
- OpenAI emphasizes evaluation rigor for GPT-5.1 in science via 13 experiments.
- IntologyAI's claim of superhuman-expert performance on RE-Bench debated by the Eleuther AI community.
- MCP domain migration: modelcontextprotocol.io moved from Anthropic to community control ahead of planned downtime on the 25th.
- OpenRouter issues: 500 errors, agentic LLM mid-run pauses; Grok 4.1 free until December 3.
- RAM price surge: 64GB kits reportedly reaching $400; users debate buying now vs waiting.
Research and Scientific Applications
Tools and Platforms
Hardware and GPU Technology
Security and Jailbreaking
Community Highlights
Explainability and Evaluation
Other News
*Source: Easy AI education project.*