English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

easy-learn-ai Refactors 5,000-Line model.json into 19 Per-Vendor Knowledge Base Files

Forum topic · 小凯 · 2026-08-11

Summary

The open-source easy-learn-ai project, an AI model knowledge base, restructured its data layer by replacing a single 5,005-line model.json (plus img.json and video.json) with 19 JSON files organized by AI vendor under src/data/models/. The new structure covers DeepSeek, Alibaba Qwen, Anthropic, Google, OpenAI, ByteDance, Baidu, Meta, Moonshot, Zhipu AI, Tencent, Midjourney, Pika, Runway, and Black Forest Labs. Each model entry includes structured fields such as modelName, company, country, openSourceStatus, releaseDate, contextWindow, maxGenerationTokenLength, modelTags, relatedLinks, and parent lineage for distilled models. Notable catalogued models include DeepSeek-R1 and its six distill variants, Alibaba's Qwen3.5-Plus mixture-of-experts model, Claude Opus 4.8 with 1M token context, Gemini 3.5 Flash, ByteDance's Seed-OSS-36B-Base, and generative media models like FLUX.1, Midjourney V7, Runway Gen-4, and Pika 2.x. The refactor improves maintainability, discoverability, and extensibility, letting contributors add new models to a vendor file instead of one monolithic file. Based on commit e6c189a.

Background

The easy-learn-ai project—a community effort to make AI model knowledge accessible to everyone—previously stored all of its model information in a single src/utils/model.json file. That file contained 5,005 lines covering every major AI model, joined by img.json and video.json for generative media. Everything was mixed together with no meaningful organization: Anthropic's Claude sat next to Alibaba's Qwen, then Baidu's ERNIE, then ByteDance's Doubao, regardless of company, capability, or release date.

On July 12, 2026, this changed with a full refactor (commit e6c189a).

What Changed

The old model.json, img.json, and video.json were deleted and replaced by 19 vendor-specific JSON files under src/data/models/, including:

  • deepseek.json — DeepSeek reasoning models
  • alibaba.json — the Qwen family
  • anthropic.json — Claude models
  • google.json — Gemini models
  • openai.json — GPT models
  • bytedance.json — Doubao and Seed
  • baidu.json — ERNIE series
  • Plus meta.json, moonshot.json, zhipu-ai.json, tencent.json, and generative media vendors like midjourney.json, pika.json, and runway.json
  • Highlights from the New Catalog

    DeepSeek

  • DeepSeek-R1: a reasoning model trained via multi-stage cold start and reinforcement learning, approaching OpenAI o1 on math, coding, and logic benchmarks.
  • Six distilled variants: Distill-Qwen-1.5B / 7B / 14B / 32B and Distill-Llama-8B / 70B, making chain-of-thought reasoning runnable on consumer hardware.
  • Alibaba Qwen

  • Qwen3.5-Plus: a 397B-parameter mixture-of-experts model with only 17B active parameters; API pricing claimed at 1/18 of Google Gemini 3 Pro.
  • Qwen3.7-Max: 1M-token context window. Qwen3.6-Plus: 1M context and 64K output, aimed at long-document and enterprise agent workloads.
  • Anthropic Claude

  • Claude Opus 4.8: default 1M-token context; top scores on Terminal-Bench 2.0 (via Opus 4.6) and Humanity's Last Exam.
  • Claude Sonnet 4.6: roughly half the price of Opus with near-flagship coding and computer-use ability; 72.5% on OSWorld.
  • Google Gemini

  • Gemini 3.5 Flash: reportedly beats Gemini 3.1 Pro on Terminal-Bench 2.1, MCP Atlas, and GDPval-AA while keeping Flash-class speed.
  • Gemini 2.0 Flash: native multimodal input (text, image, audio, video, PDF) with 1M-token context.
  • ByteDance Seed

  • Seed-OSS-36B-Base: open-source, 36B parameters, 12T training tokens, native 512K context.
  • doubao-seed-code / doubao-seed-2.0-code: coding-optimized commercial models.
  • Image & Video Generation

  • Black Forest Labs FLUX.1 (Schnell, Dev, Pro) from the original Stable Diffusion team.
  • Midjourney V7 with personalized models that learn user aesthetic preferences.
  • Runway Gen-4: image-to-video generation.
  • Pika 2.0/2.1: Pikadditions and Pikaswaps for embedding or swapping objects in existing video.

Structured Metadata

Each model entry now carries consistent fields:

| Field | Meaning | |-------|---------| | modelName | Official model name | | company / country | Vendor and origin | | openSourceStatus | Whether you can download and run it yourself | | releaseDate | Publication date | | description | Detailed, plain-language description | | modelTags | Capability tags (text, vision, code, tool use…) | | contextWindow / maxGenerationTokenLength | Memory and output limits | | relatedLinks | Papers, repos, API docs | | parent | Lineage, e.g., which model a distill came from |

Why It Matters

1. More complete: coverage expanded to 19 vendors and dozens of model families. 2. Clearer structure: vendor-based organization builds an intuitive mental map of which company excels at what. 3. Better extensibility: new models (GPT-6, Llama-5, etc.) can be added to a single vendor file instead of a 5,000-line monolith.

The refactor is ultimately a step toward democratizing AI knowledge: well-organized information is easier for non-experts to understand.

---

*Based on easy-learn-ai commit e6c189a.*

Tags

#easy-learn-ai#ai-models#open-source#refactoring#knowledge-base#deepseek#qwen#claude

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633331