English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Easy AI Daily Digest | February 3, 2026: Codex App, Step-3.5-Flash, Kimi K2.5, and Agent Security

Forum topic · 小凯 · 2026-03-27

Summary

Easy AI Daily for February 3, 2026 rounds up the day's AI news across products, models, agents, infrastructure, research, and industry. OpenAI launched a standalone macOS Codex App with multi-agent parallelism and reusable Skills; Windsurf added Arena mode for side-by-side model comparison; and LM Studio 0.4.1 gained an Anthropic-compatible local endpoint. New models include StepFun's Step-3.5-Flash (196B MoE, ~11B active), Moonshot's Kimi K2.5 topping the open-source Code Arena, Falcon-H1-Tiny sub-100M micro-models, and rumored Claude Sonnet 5 sightings on Vertex AI. Infrastructure highlights: research showing 1M-token contexts can consume ~900GB of KV-cache memory, FlashAttention v3 support for AMD RDNA GPUs, and MIT's heat-powered matrix-multiplication chip. On safety, a red-team audit of OpenClaw exposed memory-file injection risks. Industry moves: Waymo reportedly raising $16B at a ~$110B valuation, and xAI's Grok Imagine generating 1.2B videos in a month.

Easy AI Daily | February 3, 2026

A daily roundup of AI industry news curated by the Easy AI community forum.

Products & Applications

  • OpenAI releases standalone Codex App for macOS: No longer a VSCode plugin — it integrates multi-agent parallel execution, per-task git worktrees, a /plan planning mode, reusable Skills, and scheduled Automations into a single "command center" for developers.
  • Links: Official intro | Codex product page | X announcement
  • Windsurf launches Arena mode (Wave 14): Compare multiple models side-by-side on the same coding task; Battle Groups set to zero credit cost for a week, plus personal and public leaderboards.
  • Link: Windsurf download
  • LM Studio 0.4.1 adds Anthropic protocol support: A local /v1/messages endpoint lets Claude Code and similar tools swap backends to local GGUF/MLX models by just changing the base URL; OpenAI-compatible endpoint and TypeScript SDK also available.
  • Link: LM Studio blog
  • PerpetualBooster v1.1.2: Rust-based gradient boosting that replaces hyperparameter tuning with a single "budget" parameter — claims ~100x speedup over LightGBM+Optuna with similar accuracy; supports ONNX export and saving as XGBoost.
  • Link: GitHub
  • BCG deploys 36,000+ custom GPTs for 32,000 consultants, fine-tuned by role and methodology with project memory and team sharing — described as internal SaaS infrastructure.
  • AEGIS-FLOW: Community multi-agent cloud security framework (LangGraph + MCP + Next.js) that scans AWS configs, generates Terraform remediation patches, and requires human approval before applying changes. Demo
  • Lutum Veritas: A solo-built deep search engine claiming deeper research than ChatGPT/Gemini/Perplexity at ~$0.20/query, with mandatory citations, BYO API keys, and an ASK mode for verification. GitHub
  • Models & Capabilities

  • StepFun Step-3.5-Flash: 196B-parameter sparse MoE with ~11B active, aimed at long-context and agentic coding. Reported scores: SWE-bench Verified 74.4%, Terminal-Bench 2.0 51%. Day-zero vLLM support, plus Int4 quantization running 256K context on 128GB machines.
  • Links: Announcement | FP weights | Int4 weights
  • Moonshot Kimi K2.5: Ranked #1 open-source model (5th overall) in LMArena's Code Arena; Perplexity has integrated K2.5 into Pro/Max tiers.
  • GLM-4.7 Flash: Developers praise its frontend/interactive web coding; a popular cost-effective stack pairs GLM for execution with Claude/Kimi for review.
  • Claude Sonnet 5 rumors: Vertex AI logs show claude-sonnet-5 404s; rumors cite 1M context, ~50% cheaper than Opus 4.5, TPU-optimized throughput, and 80.9% on SWE-Bench. The community remains skeptical.
  • Falcon-H1-Tiny (TII): Sub-100M specialized micro-models using "anti-curriculum" training, the Muon optimizer, and hybrid Mamba+Attention. The 90M tool-calling model hits 94% relevance detection; a 600M reasoning variant scores 75% on AIME24.
  • Assistant_Pepe_8B: Fine-tuned from Nemotron on expanded 4chan data, reportedly beating its base on benchmarks — sparking debate about alignment tax and unusual statistical properties of the corpus.
  • Agents & Tooling

  • Best practice: test-first bug fixing in CLAUDE.md/AGENTS.md — write a reproduction test, fix, then prove it with the test. Widely cited as the single most effective prompt for coding agents. Thread
  • "Conductor-style engineering": One developer orchestrating 5–10 parallel agents; critics warn that context-switching overhead degrades quality.
  • OpenClaw ecosystem risks: Fun for autonomous agents, but a security audit scored it 2/100; memory files proved injectable, and OpenRouter credits can burn fast.
  • LangChain deepagents: New JS package distilling four proven agent architecture patterns (à la Claude Code, Manus) with observability and evals.
  • Recursive Language Models (RLM) in practice: Community demos using RLMs for code security auditing; tool-calling and Deno sandbox permissions remain fiddly in DSPy.
  • Infrastructure & Hardware

  • Long-context inference is memory-bound: Imperial College + Microsoft survey finds a 1M-context, batch-size-1 DeepSeek-R1 request can need ~900GB of KV-cache memory; heterogeneous prefill/decode architectures proposed. dair.ai thread
  • FlashAttention v3 lands on AMD RDNA GPUs via merged PR. PR #2178
  • Triton-Viz 3.0: Visualizes every load/store/matmul with OOB detection and inefficient-loop profiling; supports Triton and Amazon NKI.
  • Blackwell (sm120) benchmarks: TMA+mbarrier slightly beats cp.async on large matrices, but cuBLAS still appears to use sm80-era kernels.
  • MIT heat-powered chip: Performs matrix-vector multiplication using on-chip thermal gradients; currently limited to 2×2/3×3 matrices at ~99% accuracy.
  • Fudan "sushi-roll" flexible fiber chip (Nature): One meter of hair-thin fiber integrates CPU-class transistor density, withstanding 15.6 tons of pressure, repeated bending, and 100°C.
  • Research & Methods

  • Why coding agents work: Verifiable domains (compilers, tests) plus rich symbolic toolboxes equal a neurosymbolic setup; replicating it elsewhere requires building equivalent tool+verification layers.
  • Synthetic pretraining and RLVR: New method converts ordinary web text into RLVR-style reasoning tasks by masking reasoning steps and generating distractors — reviving models saturated on existing RLVR data.
  • Perplexity skepticism: Researchers warn next-token prediction quality doesn't guarantee improvements in instruction following, tool-calling stability, or multi-turn consistency.
  • ConceptMoE: Clusters similar tokens into "concept units" before MoE routing to cut redundant computation on long inputs.
  • Token-level data filtering: Work with Alec Radford's involvement proposes fine-grained token-level weighting during pretraining instead of wholesale dataset swaps.
  • Brains vs. LLMs (Nature): Human speech-processing temporal hierarchies align with LLM layer depth — deeper layers correspond to later, higher-order language cortices. Paper
  • Policy, Governance & Security

  • OpenClaw red-team exercise: Attackers bypassed defenses via shell-expansion variables embedded in JSON metadata after social engineering and pipeline RCE failed. Long-term memory stored in .md files is the biggest attack surface; credential isolation and blast-radius control are essential. Full report
  • Prompt-injection defense combo: Embedding-similarity filtering plus Grammar Constrained Decoding — filter malicious inputs, then structurally constrain what the model can output. Exercise site
  • Industry & Business

  • Waymo reportedly raising $16B at ~$110B valuation, with at least $13B from Alphabet and participation from Sequoia, DST, and Dragoneer — up sharply from $45B in October 2024.
  • xAI Grok Imagine 1.0: Generates 10-second 720p video with audio; over 1.2 billion videos created in the past 30 days.
---

📌 Source: Easy AI Daily

Tags

#ai-news#daily-digest#openai#codex#llm#ai-agents#kimi-k2-5#ai-hardware

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169294