Easy AI Daily | February 3, 2026
A daily roundup of AI industry news curated by the Easy AI community forum.
Products & Applications
- OpenAI releases standalone Codex App for macOS: No longer a VSCode plugin — it integrates multi-agent parallel execution, per-task git worktrees, a /plan planning mode, reusable Skills, and scheduled Automations into a single "command center" for developers.
- Links: Official intro | Codex product page | X announcement
- Windsurf launches Arena mode (Wave 14): Compare multiple models side-by-side on the same coding task; Battle Groups set to zero credit cost for a week, plus personal and public leaderboards.
- Link: Windsurf download
- LM Studio 0.4.1 adds Anthropic protocol support: A local
/v1/messagesendpoint lets Claude Code and similar tools swap backends to local GGUF/MLX models by just changing the base URL; OpenAI-compatible endpoint and TypeScript SDK also available. - Link: LM Studio blog
- PerpetualBooster v1.1.2: Rust-based gradient boosting that replaces hyperparameter tuning with a single "budget" parameter — claims ~100x speedup over LightGBM+Optuna with similar accuracy; supports ONNX export and saving as XGBoost.
- Link: GitHub
- BCG deploys 36,000+ custom GPTs for 32,000 consultants, fine-tuned by role and methodology with project memory and team sharing — described as internal SaaS infrastructure.
- AEGIS-FLOW: Community multi-agent cloud security framework (LangGraph + MCP + Next.js) that scans AWS configs, generates Terraform remediation patches, and requires human approval before applying changes. Demo
- Lutum Veritas: A solo-built deep search engine claiming deeper research than ChatGPT/Gemini/Perplexity at ~$0.20/query, with mandatory citations, BYO API keys, and an ASK mode for verification. GitHub
- StepFun Step-3.5-Flash: 196B-parameter sparse MoE with ~11B active, aimed at long-context and agentic coding. Reported scores: SWE-bench Verified 74.4%, Terminal-Bench 2.0 51%. Day-zero vLLM support, plus Int4 quantization running 256K context on 128GB machines.
- Links: Announcement | FP weights | Int4 weights
- Moonshot Kimi K2.5: Ranked #1 open-source model (5th overall) in LMArena's Code Arena; Perplexity has integrated K2.5 into Pro/Max tiers.
- GLM-4.7 Flash: Developers praise its frontend/interactive web coding; a popular cost-effective stack pairs GLM for execution with Claude/Kimi for review.
- Claude Sonnet 5 rumors: Vertex AI logs show
claude-sonnet-5404s; rumors cite 1M context, ~50% cheaper than Opus 4.5, TPU-optimized throughput, and 80.9% on SWE-Bench. The community remains skeptical. - Falcon-H1-Tiny (TII): Sub-100M specialized micro-models using "anti-curriculum" training, the Muon optimizer, and hybrid Mamba+Attention. The 90M tool-calling model hits 94% relevance detection; a 600M reasoning variant scores 75% on AIME24.
- Assistant_Pepe_8B: Fine-tuned from Nemotron on expanded 4chan data, reportedly beating its base on benchmarks — sparking debate about alignment tax and unusual statistical properties of the corpus.
- Best practice: test-first bug fixing in CLAUDE.md/AGENTS.md — write a reproduction test, fix, then prove it with the test. Widely cited as the single most effective prompt for coding agents. Thread
- "Conductor-style engineering": One developer orchestrating 5–10 parallel agents; critics warn that context-switching overhead degrades quality.
- OpenClaw ecosystem risks: Fun for autonomous agents, but a security audit scored it 2/100; memory files proved injectable, and OpenRouter credits can burn fast.
- LangChain deepagents: New JS package distilling four proven agent architecture patterns (à la Claude Code, Manus) with observability and evals.
- Recursive Language Models (RLM) in practice: Community demos using RLMs for code security auditing; tool-calling and Deno sandbox permissions remain fiddly in DSPy.
- Long-context inference is memory-bound: Imperial College + Microsoft survey finds a 1M-context, batch-size-1 DeepSeek-R1 request can need ~900GB of KV-cache memory; heterogeneous prefill/decode architectures proposed. dair.ai thread
- FlashAttention v3 lands on AMD RDNA GPUs via merged PR. PR #2178
- Triton-Viz 3.0: Visualizes every load/store/matmul with OOB detection and inefficient-loop profiling; supports Triton and Amazon NKI.
- Blackwell (sm120) benchmarks: TMA+mbarrier slightly beats cp.async on large matrices, but cuBLAS still appears to use sm80-era kernels.
- MIT heat-powered chip: Performs matrix-vector multiplication using on-chip thermal gradients; currently limited to 2×2/3×3 matrices at ~99% accuracy.
- Fudan "sushi-roll" flexible fiber chip (Nature): One meter of hair-thin fiber integrates CPU-class transistor density, withstanding 15.6 tons of pressure, repeated bending, and 100°C.
- Why coding agents work: Verifiable domains (compilers, tests) plus rich symbolic toolboxes equal a neurosymbolic setup; replicating it elsewhere requires building equivalent tool+verification layers.
- Synthetic pretraining and RLVR: New method converts ordinary web text into RLVR-style reasoning tasks by masking reasoning steps and generating distractors — reviving models saturated on existing RLVR data.
- Perplexity skepticism: Researchers warn next-token prediction quality doesn't guarantee improvements in instruction following, tool-calling stability, or multi-turn consistency.
- ConceptMoE: Clusters similar tokens into "concept units" before MoE routing to cut redundant computation on long inputs.
- Token-level data filtering: Work with Alec Radford's involvement proposes fine-grained token-level weighting during pretraining instead of wholesale dataset swaps.
- Brains vs. LLMs (Nature): Human speech-processing temporal hierarchies align with LLM layer depth — deeper layers correspond to later, higher-order language cortices. Paper
- OpenClaw red-team exercise: Attackers bypassed defenses via shell-expansion variables embedded in JSON metadata after social engineering and pipeline RCE failed. Long-term memory stored in .md files is the biggest attack surface; credential isolation and blast-radius control are essential. Full report
- Prompt-injection defense combo: Embedding-similarity filtering plus Grammar Constrained Decoding — filter malicious inputs, then structurally constrain what the model can output. Exercise site
- Waymo reportedly raising $16B at ~$110B valuation, with at least $13B from Alphabet and participation from Sequoia, DST, and Dragoneer — up sharply from $45B in October 2024.
- xAI Grok Imagine 1.0: Generates 10-second 720p video with audio; over 1.2 billion videos created in the past 30 days.
Models & Capabilities
Agents & Tooling
Infrastructure & Hardware
Research & Methods
Policy, Governance & Security
Industry & Business
📌 Source: Easy AI Daily