English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | November 20, 2025: Gemini 3 Pro Image, Olmo 3, SAM3, GPT-5.1 Codex Max and More

Forum topic · 小凯 · 2026-03-27

Summary

A comprehensive daily digest of AI industry news for November 20, 2025. Google released Gemini 3 Pro Image (Nano Banana Pro) with Google Search grounding, 2-4K resolution, and text rendering error rates cut from 56% to 8%. AI2 fully open-sourced Olmo 3 under Apache-2.0, including a 32B Think reasoning variant. Meta launched SAM3 and SAM3D segmentation models with unified image/video segmentation, 30ms inference, and single-image 3D reconstruction. OpenAI introduced GPT-5.1 Codex Max with native multi-context window support and SOTA on SWEBench, plus shared 13 early experiments using GPT-5.1 for scientific research. Perplexity launched its Comet browser on Android, Mac, and Windows; Cursor rolled out beta debug mode; and community discussions covered Jetson cluster builds, CUDA/DMA optimizations, SynthID watermark bypasses, jailbreaking concerns, and a sharp rise in RAM prices (64GB kits reaching $400).

📅 AI Industry Highlights — November 20, 2025

This digest is an English translation of the Easy AI Daily report, originally published on zhichai.net.

Key points

Model Updates and Releases

Google releases Gemini 3 Pro Image (Nano Banana Pro)

  • Supports Google Search grounding, 2–4K resolution, and text-in-image generation/editing
  • Pricing: $0.134 per 2K image, $0.24 per 4K image
  • Available via Gemini App/API, LM Arena, Hugging Face Spaces, and Together AI
  • Early demos show accurate infographics and chart annotation; text rendering error rate dropped from 56% to 8%
  • Links: Pricing | Announcement | LM Arena | Hugging Face Spaces | Together AI | Error rate data | SynthID watermarking
  • AI2 releases Olmo 3 fully open models

  • Fully open source under Apache-2.0, including a 32B Think variant for long chain-of-thought and complex reasoning
  • Post-norm architecture retained; 7B uses sliding-window attention for KV cache optimization, 32B uses GQA
  • RL infrastructure delivers 4x faster experiments; emphasis on decontaminated evaluation (e.g., random reward tests)
  • Links: Announcement | Architecture analysis | Hugging Face listing
  • Meta releases SAM3 and SAM3D segmentation models

  • SAM3 unifies image/video segmentation with text/visual prompts, 2x performance improvement, 30ms inference
  • SAM3D enables 3D reconstruction from a single image
  • Data engine: 4M phrases, 52M masks; source code released for commercial use
  • Links: SAM3 announcement | SAM3D announcement | GitHub repo
  • OpenAI launches GPT-5.1 Codex Max

  • Designed for long-running, detail-heavy tasks; first native multi-context-window support (via compaction)
  • SOTA on SWEBench; available only through ChatGPT plans, no API access
  • Link: Release blog
  • Cogito 2.1 enters WebDev Arena

  • Deep Cogito's model ranks #18 overall, top-10 among open models; hosted on Together and Fireworks
  • Link: Model page
  • Research and Science

  • OpenAI shares 13 early experiments using GPT-5.1 for scientific research, covering math, physics, biology, and materials science; four experiments helped solve previously unsolved problems. Includes blog, technical report, and researcher podcast.
  • Links: Overview | arXiv thread
  • Tools and Platforms

  • Perplexity Comet browser launched for Android, Mac, and Windows: voice-first browsing, support for Kimi-K2 Thinking and Gemini 3 Pro; Pro/Max users can create slides, tables, and documents
  • Cursor beta debug mode: new log-ingest server auto-instruments code to collect logs; agents validate hypotheses from logs rather than guessing
  • MemMachine Playground open on Hugging Face Spaces: persistent AI memory with GPT-5, Claude 4.5, Gemini 3 Pro; fully open source
  • DSPy Proxy repo (aryaminus/dspy-proxy) released, simplifying DSPy agent development: GitHub
  • Hardware and GPU Tech

  • A user built a cluster from six NVIDIA Jetson devices for NCCL/NVIDIA development, testing pre-B300-cluster workflows
  • GPU MODE discussions: GEMM optimization, CUDA texture vs. constant cache, AMD MI300X DMA collectives (+16% for large transfers, paper), BF16 conversion issues
  • Mojo 0.25.7 nightly shows a major regression on Mac M1: llama2.mojo throughput fell from ~1000 tok/sec to ~170 tok/sec
  • Safety and Jailbreaking

  • BASI community discussed jailbreaks against Gemini 3 Pro, Grok (shell access obtained), and Claude 4.5 (trust-building bypass)
  • Users reported bypassing Gemini's SynthID watermark via a "do nothing" edit prompt on reve-edit, or by directly asking the model if an image is AI-generated
  • Community Highlights

  • LMArena: debates on Nano Banana Pro text rendering, GPT-5.1 vs. Gemini 3 Pro
  • Perplexity community: Gemini 3 Pro coding praised over Claude Sonnet 4.5; Comet RAM usage concerns
  • Unsloth AI: RAM prices surging (64GB kits at $400)
  • Yannick Kilcher community: AI CEO benchmark (LLM long-horizon planning below human), NVIDIA Q3 earnings
  • Eleuther AI: KNN vs. quadratic attention, IntologyAI's RE-Bench results claiming superhuman expert performance
  • Moonshot AI community: Kimi K2 Coding plan ($19) pricing complaints; SGLang tool-calling issues
  • Other News

  • modelcontextprotocol.io migrated to community control ahead of scheduled downtime on the 25th
  • OpenRouter users reported 500 errors and mid-task pauses in agentic LLM calls; Grok 4.1 free until December 3
  • RAM prices surged sharply, prompting buy-now-or-wait debates
---

*Source: Easy AI teaching project*

Tags

#ai-news#google-gemini-3#olmo-3#sam3#gpt-5-1-codex-max#open-source-models#ai-hardware#daily-digest

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169141