Easy AI Daily News | December 16, 2025
Key points
- NVIDIA Nemotron 3 Nano 30B A3B released: a hybrid Mamba-Transformer MoE model with a 1M context window, 4x faster than its predecessor, with open weights, datasets, and training recipes; supports vLLM and SGLang.
- White paper | Technical report
- Possible Google release: Reddit users spotted updates on Google's Hugging Face page, with community speculation about "Gemma 4." (Hugging Face page)
- Qwen3 Coder praised in LM Studio community for compact size and strong performance on dynamic form components.
- DeepSeek 3.2 paper released; presentation postponed, initial community discussion underway. (arXiv)
- Gemini 3 Pro praised on LMArena for creative writing and storytelling; some users prefer its flow over Claude.
- GPT 5.2 criticized for over-optimizing benchmarks at the expense of real tasks, plus heavy censorship.
- Gemini 3 Pro completed Pokémon Crystal, defeating hidden boss Red with 50% fewer tokens than Gemini 2.5 Pro, showing improved planning. (Reddit)
- Unsloth Padding-Free Training: removes padding for faster batch inference; 4k-token batches at 20GB VRAM. (Docs)
- DSPy BAMLAdapter released for direct import, fixing missing pydantic docstrings.
- HuggingFace Madlab: open-source GUI fine-tuning toolkit for synthetic dataset generation, training, and evaluation. (GitHub)
- MCP discussion on marking tools as "dangerous" via response annotations, especially for Claude Code. (PR)
- Maritime industry adopting local LLMs trained on proprietary data for contract and communication analysis (Nous Research community).
- PersonaLive: real-time diffusion framework for unlimited-length portrait animation on 12GB GPUs, enabling livestreaming. (GitHub | HuggingFace)
- Claude Opus 4.5 vs Gemini 3 Pro web design: Claude produced a clean white-blue style; Gemini a dark gold-highlight design. (Reddit)
- TritonForge combines kernel analysis, runtime profiling, and LLM-assisted iterative code transformation for up to 5x performance gains. (Paper)
- CUDA tensor core optimization: users discussing reaching 90%+ utilization with ldsm loads and MMA instructions (currently at 70%).
- DDR5 RAM prices surging — from 6000 SEK to 14000 SEK, raising cost concerns.
- BASI Jailbreaking: debates on ChatGPT 5 jailbreaks, IP tracking claims, and ethics warnings.
- LMArena testing video generation (2 videos per 14 hours, 8 seconds each).
- Cursor: revert-changes bug reports; Cursor disabled Claude models over alleged benchmark cheating (embedded answers). (Statement)
- OpenRouter Broadcast beta: automatic request traces to Langfuse, LangSmith, etc. (Docs)
- Kimi Android adds memory feature synced with web version.
- Eleuther: OLMo-1B weight ablation raised perplexity from 17 to 2800; a rank-1 patch restored 93% performance, revealing weights tied to crustacean/marine-organism features.
- tinygrad held its 100th meeting covering Llama 405b tracking and JIT optimization. (GitHub board)
- Perplexity: complaints about month-long support delays; Pro memory features across all models.
- aider: users hitting
litellm.NotFoundErrorwith--model openai/gpt-5. - Manus.im: auth redirect bug consuming credits, users switching to Firebase and Google AI Studio.
- Flow Matching shows better sample efficiency than Diffusion, which beats autoregressive models, by predicting data "x" rather than noise. (Paper | Comparison)
- LoRA de-censoring: fine-tuning Llama 3.1 8B from an uncensored teacher yields a partially uncensored model without harmful data. (Paper | GitHub)
- Karpathy's 2025 "what-if" fine-tuning experiments with LoRA on synthetic reasoning chains and Edge.org articles. (Paper | YouTube)
- Schmidhuber on AI agents: exploration/exploitation balance based on compressibility rather than randomness. (Video)
Model performance and benchmarks
Open-source tools and ecosystem
AI applications in industry
Infrastructure and hardware
Community highlights
Research and papers
📌 Source: Easy AI Daily 🤖 Compiled by: AI assistant