RC RANDOM CHAOS

AI infrastructure

7 posts

Zhipu built its own inference stack for GLM
Article

Zhipu built its own inference stack for GLM

GLM built its own inference infrastructure to serve LLMs cheaply on constrained, mixed hardware. Here's what breaks in generic stacks and what to copy.

Four bits, then a cliff
Article

Four bits, then a cliff

How to choose LLM quantization for production: why 4-bit is the default, where 1-bit collapses, and why your own eval set is the real deployment gate.

The demo passed. Two weeks later, the queue filled.
Article

The demo passed. Two weeks later, the queue filled.

Prompt engineering treats AI as magic. Reliable LLM systems come from validation, retries, fallbacks, and monitoring - not better wording.

Stanford teaches LLMs by making you build one
Article

Stanford teaches LLMs by making you build one

What CS336 actually teaches LLM engineers, where the course exposes silent drift, and why the skills transfer directly to RAG, agents, and eval.

Liquid AI's 8B-A1B drop rewrites inference math
Article

Liquid AI's 8B-A1B drop rewrites inference math

Liquid AI's 8B-A1B MoE trained on 38T tokens shifts LLM inference economics. What it means for engineering pipelines and workforce planning.

Hugging Face revived PapersWithCode in early 2025
Article

Hugging Face revived PapersWithCode in early 2025

Hugging Face's PapersWithCode revival restores the verification substrate LLM engineering teams lost, reshaping pipelines and AI workforce roles.

Article

AI costs more than humans

Nvidia says AI costs more than human workers. The real issue is architecture, not compute price. Here is how to fix the unit economics.