Skip to content

Mining cards, made to serve tokens

The NVIDIA CMP 100-210 / CMP 100HX-210 uses GV100 silicon and HBM2 memory, but native CMP configuration has fewer usable compute units, severely reduced tensor and FP64 throughput, and a PCIe Gen1 x1 connection. The tested multi-card system also lacks a usable peer route. The hardware comparison records the measured differences and their limits.

This site is the engineering record of making llama.cpp run well on them anyway.

  • Measure compute and transport


    Tensor throughput, host transfers and scheduling can each limit a request. Measure the intended model and device split before choosing a kernel or transport optimization.

  • Byte-for-byte or it does not ship


    A patch is retained only if enabling it leaves the generated tokens identical to leaving it off, at every context in the ladder. Faster but different is a rejected experiment, not an optimization.

  • Activation is documented


    Each result identifies the affected configurations and whether a runtime selector controls the change. Accepted source specializations can become part of the normal build.

  • Promotion includes the tradeoffs


    The public timeline contains permanently adopted patches, with their measured regressions and validation limits. Ongoing and undecided tests remain private. Publication accompanies a permanent patch promotion.

The shape of the machine

Everything in the patch series follows from one picture. Weights and activations must cross a link roughly three orders of magnitude narrower than the memory the cards themselves have, and any card-to-card traffic crosses it twice.

flowchart LR
    subgraph HOST["Host"]
        CPU["CPU + system RAM"]
        PIN["mapped pinned staging"]
    end
    subgraph FABRIC[" "]
        LINK["PCIe Gen1 x1<br/>~250 MB/s each way"]
    end
    subgraph CARDS["CMP 100-210 &times; N"]
        G0["GV100 &middot; SM70<br/>HBM2 ~850 GB/s"]
        G1["GV100 &middot; SM70<br/>HBM2 ~850 GB/s"]
    end

    CPU --- PIN
    PIN <--> LINK
    LINK <--> G0
    LINK <--> G1
    G0 -. "no P2P, no NVLink<br/>every hop stages through the host" .-> G1

    classDef slow fill:#c2410c,stroke:#7c2d12,color:#fff
    classDef fast fill:#0f766e,stroke:#134e4a,color:#fff
    class LINK slow
    class G0,G1 fast

Three consequences drive every decision on this site:

  1. Bytes on the link are the scarce resource. The right move is usually to remove a transfer, not to compress it. That is what device-side mask expansion does.
  2. Cross-card boundaries are host round trips. With no peer path, a tensor handed from card 3 to card 4 goes down the link and back up it. Making that staging explicit and reusable is the mapped pinned bridge.
  3. Submission cost is not amortised by busy GPUs. With mostly idle cards, driver submission dominates sampled CPU time, which is why graph capture matters more here than it would on a saturated node.

Native CMP identification rejected Nsight CUDA profiling in the tested configuration. Separate diagnostic captures on a V100-reporting configuration informed the D256 investigation; those kernel timings are not native-CMP performance measurements. Public results distinguish ordinary inference timing from diagnostic profiling. See Method.

The series

All eight patches are in one ordered series. Seven earlier mechanisms and the accepted D256 specialization compile together in the public source. The results timeline records accepted measurements and their limits; source reconstruction alone is not performance certification.

Patch What it changes Selector
Research selector registry Infrastructure: selectors become runtime-switchable and their reads observable
DP4A MMQ routing Forces the integer-dot MMQ route for quantized prefill --cuda-mmq force
Volta Q4_K and Q5_K DP4A Extends that routing to two quantizations SM70 did not cover --cuda-mmq force
Parallel model load Overlaps independent per-GPU uploads with positional reads LLAMA_MODEL_LOAD_PARALLEL
Mapped pinned host bridge Explicit, event-ordered staging for no-P2P card boundaries GGML_CUDA_MAPPED_HOST_BRIDGE
Prompt graph capture seed Replays prompt graphs after a safe capture-only first observation GGML_CUDA_PROMPT_GRAPH_CAPTURE_SEED
Device-side mask expansion Sends a compact visibility hint and rebuilds the mask on the GPU

Eighth patch: SM70 D256 shared Q and unit rescaling documents the additional source specialization and its measured tradeoffs.

Read the series overview See stock-versus-fork benchmarks Browse the timeline

Reading this site

  • The hardware

    What a CMP 100-210 actually is, what it is missing, and which of those absences cost the most.

  • The patches

    One page per patch: the mechanism, the diagram, the selector, and what it is not claimed to do.

  • The timeline

    Permanently promoted patches with their measured tradeoffs, newest first. Backed by a machine-readable index for agents.

  • Method

    The promotion ladder, the byte-exactness gate, and why activation is proved by a log line rather than by an environment variable being set.

Attribution

Built and maintained by Joseph Stackhouse at Stack-Tech. The patches are offered against upstream llama.cpp under its MIT licence. See About.