Easy AI Daily | December 9, 2025
*Translated from the zhichai.net Easy AI Daily digest.*
Model Updates & Multimodal Progress
- Zhipu releases GLM-4.6V and GLM-4.6V-Flash: GLM-4.6V is a 106B MoE model for cloud/high-performance clusters, while GLM-4.6V-Flash is a 9B dense model for local, low-latency use. Both support 128k context and native multimodal function calling. The Flash API is free, and weights are open on Hugging Face. (GLM-4.6V-Flash | GLM-4.6V)
- JinaAI releases Jina-VLM-2B: A 2B-parameter multilingual VLM focused on charts, documents, and scene text. It averages 72.3 across 8 VQA benchmarks, with SOTA on MMMB (78.8) and multilingual MMBench (74.3). (Announcement)
- Qwen3-4B shines locally: Reaches 70 tokens/sec on an RTX 2060 with strong coding ability, though it tends to "overthink"; users recommend pairing it with GLM-4.6V-Flash.
- DeepSeek V3.2 improves reasoning: Supports interleaved reasoning with strong RoO code performance. Pricing: $0.28/M input tokens, $0.45/M output tokens; reportedly more stable than Kimi.
- Hugging Face launches Claude Code skills for LLM training automation: Specify training tasks in natural language (e.g., fine-tuning Qwen3-0.6B); the system handles data validation, GPU selection, and HF Jobs submission. Small runs cost around $0.30. (Blog)
- Unsloth AI progress: Celebrated 10k Reddit members, fixed slow HF downloads, and released Mistral Large 3 GGUF models for local inference. (Reddit | GGUF)
- DSPy TOON Adapter: Reduces token consumption, though it handles nested schemas less well than BAMLAdapter; MMLU-Pro improves after GEPA optimization. (Code)
- LangChain Deep Agents evaluation: A framework for evaluating long-running agents (planning, filesystem, sub-agents), averaging 42.65% on Terminal Bench 2.0, with context-compression triggers added. (Announcement)
- Google's Miras post-Transformer framework: A NeurIPS paper frames Transformers/RNNs as associative memory systems, outperforming Transformer/Mamba2/DeltaNet on LM, reasoning, and long-context tasks, with a 20% improvement in long-text retrieval. (Overview)
- AxiomProver solves 9/12 Putnam problems: The Lean-based system solved 9 problems within hours of the exam, emphasizing verifiability and formal pipelines, exceeding last year's top performance. (Announcement)
- NeurIPS mechanistic interpretability workshop: Chris Olah reflected on interpretability, arguing the field needs scalable tools rather than neuron-level analysis of single models.
- MEMTRACK benchmark: Evaluates agent long-term memory via Slack/Linear/git scenarios; GPT-5 scored 60%, showing room for improvement in real tool environments. (Announcement)
- Dexter 2.0: An open-source financial research agent with planning and self-verification, built on LangChain, suited for long-horizon financial analysis. (Demo)
- AI21 Maestro: Agent orchestration with multi-step planning, built-in verification, proprietary RAG, and execution graphs. (Announcement)
- Cursor Agent issues: Community reports of failures creating files and infinite loops; approval-button workarounds exist but a permanent fix is needed.
- OpenAI Stargate triggers DRAM shortage: Stargate will consume up to 40% of global DRAM output (~900k wafers/month) via deals with Samsung and SK Hynix, driving up DDR5 prices and affecting even gamers' memory costs. (Tom's Hardware)
- AMD 7900XTX as budget AI GPU: Great price/performance with solid llama.cpp support for local LLMs; ray tracing lags but AI performance is strong. (Discussion)
- Blackwell WGMMA incompatibility: Blackwell GPUs lack WGMMA (sm90a-only), requiring kernel rewrites; CUDA 13.1 introduces CUDA Tile to simplify programming. (Blog)
- RadixArk spins off from SGLang: Focused on AI scheduling, compilers, serving, and training pipelines to make frontier infrastructure more open. (Announcement)
- Local LLM cluster build: One user assembled 8x RTX 3090 (192GB VRAM) + 64-core EPYC Milan + 250GB RAM (~$8k), running GLM-4.5 Air Q6_K at 49 tokens/sec via llama.cpp. (Details)
- Vector database selection guide: HNSW (<10M vectors), Turbopuffer (large databases), pgvector (small/local), Chroma (lightweight); the guide criticizes some commercial DBs for inefficiency. (Blog)
- Core community contributors: Unsloth (fast fine-tuning), mradermacher (automated quantization), Bartowski (curated quants), and TheBloke (base models) recognized by users. (Discussion)
- picomon AMD GPU monitor: Open-source AMD GPU monitoring from a HF user, more reliable than nvtop at the cost of some precision. (Code)
- Meta acquires Limitless: Formerly Rewind; Pendant users get 1 year of support, while non-Pendant features will sunset. (Announcement)
- ARC Prize 2025 results: NVARC took Top Score at 25.03%; the TRM paper won $50k; the $600k Grand Prize went unclaimed. Winning methods will be open-sourced. (Results)
- Manus.im billing issues: Users report lost credits and auto-renewing subscriptions to 2026, with slow support responses, suspected system bug.
- Sora 2 regional restrictions: Available in only 7 countries; VPN use violates the ToS and may lead to account bans. (Supported countries)
- ChatGPT "napping on the job" screenshots go viral. (Meme)
- The "AI fixes a bug" loop meme captures developer frustration with AI claiming fixes that don't work. (Video)
- Ultra-detailed David Duchovny AI image with a 20-item checklist verification. (Tweet)
- Nano Banana Pro fashion editorial workflow using contact-sheet prompts with camera and styling constraints. (Details)
Training & Optimization Tools
Research & Evaluation
Agents & Workflows
Infrastructure & Hardware
Community & Ecosystem
Industry & Market
Fun & Memes
*Source: Easy AI education project*