HELIX Columnar Embedding — Bandwidth Report
Generated: 2026-05-19T05:33:18Z
Model
- Model:
Qwen/Qwen2.5-0.5B-Instruct - Vocab size: 151,936
- Hidden dim: 896
- Embedding dtype:
float32 - Row bytes: 3,584 B (896 × 4 B)
- Full embedding matrix: 519.31 MB
Per-session bandwidth
| Workload | Prompts | Total tokens | Unique rows | Columnar bytes | Dense ship | Ratio | Saved |
|---|---|---|---|---|---|---|---|
qa_short |
5 | 33 | 27 | 94.50 KB | 519.31 MB | 0.0178% | 99.9822% |
qa_medium |
3 | 68 | 53 | 185.50 KB | 519.31 MB | 0.0349% | 99.9651% |
code_python |
3 | 81 | 49 | 171.50 KB | 519.31 MB | 0.0323% | 99.9677% |
chat_turn |
1 | 66 | 53 | 185.50 KB | 519.31 MB | 0.0349% | 99.9651% |
long_context |
1 | 628 | 109 | 381.50 KB | 519.31 MB | 0.0717% | 99.9283% |
multilingual |
5 | 38 | 31 | 108.50 KB | 519.31 MB | 0.0204% | 99.9796% |
Multi-session amortization
Simulating 100 independent sessions on one node with top-100 popular tokens pre-warmed in the row cache. (Realistic for a multi-tenant edge node where BOS/EOS/common BPE pieces saturate quickly.)
| Metric | Value |
|---|---|
| Sessions | 100 |
| Warm-set size | 100 rows |
| Warm-seed bytes (one-time) | 350.00 KB |
| Cold rows fetched (total) | 185 |
| Cold bytes fetched (total) | 647.50 KB |
| Columnar total (warm-seed + cold) | 997.50 KB |
| Dense ship per session × N | 50.71 GB |
| Ratio | 0.001876% |
| Bandwidth saved | 99.998124% |
Notes
- Rows are exact (no spectral approximation): columnar fetch is a zero-quality-loss alternative to lossy GKQ rank-K SVD for the embedding tensor.
- These numbers are for the input embedding half. With tied embeddings (Qwen-style) the same matrix is the lm_head; see the design doc, §5, for the speculative top-k strategy that extends the win to the output projection.
- Multi-session amortization assumes one node serves all sessions. Sharded across N nodes, each node's warm-set + cold-rows scales with its session subset; the ratio against dense-ship-per-node stays comparable.