English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Running a 753B-Parameter GLM-5.2 Locally on Two Mac Studios: An Extreme Experiment

Forum topic · 小凯 · 2026-07-12

Summary

A June 2026 community experiment demonstrated that GLM-5.2, a 753-billion-parameter large language model, could run locally on two Mac Studio machines with M5 Max chips, achieving roughly 16 tokens per second. The trick was aggressive IQ1_S 1-bit quantization, which shrank the model's storage footprint to about 94GB — small enough to fit in the combined unified memory of two Macs. Notably, community benchmarks showed that the quantized GLM-5.2 outperformed a much smaller Qwen 27B model (8-bit) on some coding tasks, suggesting that scale can partially compensate for reduced precision. The same period saw leaked specs for Meituan's LongCat 2.0 / Owl Alpha: 1.6T total parameters, 48B active parameters via MoE, a 1M-token context window, reportedly trained on 50,000 domestic Chinese accelerator cards — a signal of China's shift toward a self-reliant AI compute supply chain. Together, these developments highlight two converging trends: democratization of frontier AI via local deployment, and independent domestic AI infrastructure.

A Crazy Experiment

In late June 2026, a rumor spread through the AI community that sounded like bragging: someone got GLM-5.2 753B — a 753-billion-parameter large language model — running locally on two Mac Studio machines with M5 Max chips. Not through a cloud API, not on rented servers — right on a desk, plugged into the wall, fans whirring.

The speed was 16 tokens per second. Not slow, not fast either. But if you understand what this means, it feels like magic.

For comparison: in 2023, GPT-4 was estimated at around 1.8 trillion parameters, but it was only accessible through OpenAI's API — online, registered, paid, rate-limited. Back then, saying "I want to run a near-GPT-4 model at home" would have gotten you laughed at, because GPT-4-class models required far more memory than any consumer computer had.

Three years later, a 753B model runs on two Macs. The progress involves better chips, quantization breakthroughs, and architectural optimization — and something deeper: AI is shifting from a "cloud privilege" to a "local right."

What Is "Quantization" and Why It Matters

LLM parameters are numbers, typically stored as 32-bit floats (FP32) during training. A 753B model at FP32 would need roughly 3TB of memory — far beyond consumer hardware.

Quantization asks: do we really need that much precision? Compressing 32-bit to 16-bit, 8-bit, or even 1-bit often works surprisingly well. LLMs are highly fault-tolerant; dropping precision may cost a little quality but saves enormous space.

This experiment used IQ1_S, an extremely aggressive 1-bit format. The 753B model needed only about 94GB of storage — just fitting into two Mac Studios' unified memory (48–64GB each).

The cost is precision loss. But here's the interesting part: in some tasks, a huge low-precision model can beat a small high-precision one.

Big and Coarse vs. Small and Precise

Community benchmarks compared GLM-5.2 753B (1-bit) against Qwen 27B (8-bit). Interestingly, by total information content GLM-5.2 actually holds more data (753B × 1bit = 753Gbit vs. 27B × 8bit = 216Gbit).

Result? On some coding tasks, GLM-5.2 won.

This reveals a deep rule: a model's knowledge capacity depends not only on precision but on scale. A large, coarse model may have memorized more patterns, edge cases, and programming paradigms — like someone who read 1,000 books roughly versus someone who read 10 books in detail. For tasks requiring broad knowledge, the former can win. Precision still matters for exact mathematical reasoning, but the optimal scale-precision tradeoff differs by task — opening huge flexibility for local deployment.

Meituan LongCat 2.0: Domestic Hardware Bets

The same day brought an even bigger leak: Meituan's LongCat 2.0 / Owl Alpha specs — 1.6T total parameters, 48B active, 1M context window, trained on 50,000 domestic Chinese accelerator cards.

48B active parameters indicates a MoE (Mixture-of-Experts) architecture — activating only a fraction of parameters per inference. A 1M context window handles roughly 1.5 million Chinese characters — an entire book in one go.

The standout number: 50,000 domestic accelerators. This signals:

1. Domestic chips can now support frontier-scale training. This is commercial-grade, not experimental. 2. China is building a "de-Americanized" AI compute supply chain — from chip design to data centers, a self-controlled stack is forming, as NVIDIA A100/H100 access is restricted by export controls.

These details come from community leaks, not officially confirmed by Meituan — but even as a semi-official signal, they're enough to reshape assessments of domestic AI hardware competitiveness.

Why "Running at Home" Matters

Why run models locally when cloud APIs are convenient?

  • Privacy. Data never leaves your device — medical records, legal documents, trade secrets.
  • Control. Cloud APIs can be throttled, repriced, or shut down. A downloaded model is yours, like owning a book versus borrowing one.
  • Latency. No network round trips — essential for coding assistance, real-time translation, game NPC dialogue.
  • Cost. High upfront hardware cost, but potentially cheaper than ongoing API fees for heavy use.
  • Customization. Local models can be fine-tuned to your code style, writing habits, or internal company knowledge — impossible with cloud APIs.
Local deployment has costs too: expensive hardware, complex setup, maintenance. But quantization and consumer chip advances are rapidly lowering these barriers.

Final Thoughts

GLM-5.2's local experiment and Meituan's domestic-hardware training are two sides of the same coin: democratization of AI capability — ordinary people running models once reserved for big companies — and autonomy of AI infrastructure — China building a compute base independent of external supply chains.

Together they sketch a future where large models are no longer cloud oracles but standard equipment, like operating systems, running on your computer, phone, or smart home devices.

We're not there yet: 16 tokens/s is too slow for real-time conversation, 1-bit quantization still loses too much quality on some tasks, and domestic chip ecosystems need time to mature. But the door is open. What seemed impossible three years ago now happens on someone's desk.

That's the charm of technological progress — it always arrives faster than expected.

---

*Original hashtags: easy-learn-ai, daily updates, GLM, local LLMs, domestic chips, quantization.*

Tags

#glm-5.2#local-llm#quantization#mac-studio#moe#domestic-chips#longcat-2.0#large-language-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379406