Machine learning systems

mxinO

Notes and visual explanations connecting research ideas, model architecture, and implementation details.

06 published visual explainers

Visual explainers

Standalone, browser-based guides with diagrams, equations, and interactive controls.

Preview of the DeepSeek-V4.1-Flash architecture explainer

Model architecture

DeepSeek-V4.1-Flash

Follow CED prefill and decode, CSA2 cache sharing, hierarchical sparse indexing, and bounded replay through the 40-layer model.

Open explainer
Preview of the FlashAttention explainer

GPU algorithms

FlashAttention

See how tiling and online softmax reduce HBM traffic without approximating attention.

Open explainer
Preview of the Sol-Attn explainer

Attention systems

Sol-Attn

Trace sequence parallelism, communication costs, and the system design behind long-context attention.

Open explainer
Preview of the MiniMax H3 architecture explainer

Model architecture

MiniMax H3

Explore lightning attention, mixture-of-experts routing, and the model's hybrid recurrent structure.

Open explainer
Preview of the Looped World Models explainer

World models

Looped World Models

Understand iterative latent refinement, adaptive computation, and the evidence behind recurrent depth.

Open explainer
Preview of the Blackwell GPU and FP4 explainer

Numerical formats

Blackwell GPU and FP4

Compare low-precision formats and follow FP4 values through Blackwell tensor-core execution.

Open explainer