Models and Benchmarks
Terminal-Bench 2.0 Released with Harbor Framework
Terminal-Bench 2.0 fixes issues with tasks being too easy or too hard, adopts the Harbor framework for cloud container execution, and held a launch party with a recorded video.> Links: Terminal-Bench 2.0 announcement | Launch party video
Moonshot AI Releases Kimi K2 Thinking
Kimi K2 Thinking is a 1T-parameter MoE model (32B active parameters), INT4 quantized, with 256K context, achieving an AI Index score of 67. It supports deployment via MLX and Ollama and integrates the slime framework.> Links: Kimi K2 Thinking model page | MLX deployment PR | Ollama support
Community Updates
AI Twitter Recap
Discussions on Kimi K2 performance, MoE inference optimization, long-context information aggregation, plus DreamGym synthetic environments and EdgeTAM real-time tracking tools.AI Reddit Recap
Includes Kimi K2 performance debates, discussions on AI consciousness development, free ChatGPT Go and Gemini Ultra services in India, and an AI-designed cookie box failure case.> Links: Kimi K2 creative writing discussion | Free AI services in India
AI Discord Recap
LMArena discussed Gemini 3 Pro performance, Perplexity AI discussed Kimi K2, GPU MODE covered an FP4 hackathon and Blackwell bandwidth, and OpenRouter released Embeddings and a TypeScript SDK.> Links: LMArena Discord | Perplexity AI Discord
Tools and Frameworks
Unsloth Adds MoE Fine-Tuning Support
Unsloth's FastModel tool now supports fine-tuning MoE models, addressing poor Transformers MoE support, and is compatible with both dense and sparse models.> Link: Unsloth docs
Mojo Updates and Performance Optimization
Mojo's try-except error handling outperforms Rust; CPU multithreading is not yet supported, and the compiler remains based on C++ and MLIR.DSPy FastWorkflow Achieves SOTA on Tau Bench
DSPy's FastWorkflow achieves SOTA on Tau Bench retail and airline workflows, highlighting the impact of context engineering on small models.> Link: FastWorkflow repo
Intel Releases llm-scaler
Intel's llm-scaler tool optimizes LLM performance on Intel GPUs, supports ERP models, and improves inference efficiency.> Link: llm-scaler repo
Events and Workshops
AI Scholars AI Agent Workshop
AI Scholars hosted an online and in-person workshop teaching how to build AI agents with LangChain, AgentKit, and AutoGen, based on real customer data analysis problems.> Link: RSVP
---
*Source: Easy AI teaching project*