forgo.cloud
Sign in
Repo workspace

forkjoin-ai/gnosis

HELIX Columnar Embedding — Bandwidth Report

distributed-inference/COLUMNAR_BANDWIDTH_REPORT.md
forkjoin-ai/gnosis

HELIX Columnar Embedding — Bandwidth Report

Generated: 2026-05-19T05:33:18Z

Model

  • Model: Qwen/Qwen2.5-0.5B-Instruct
  • Vocab size: 151,936
  • Hidden dim: 896
  • Embedding dtype: float32
  • Row bytes: 3,584 B (896 × 4 B)
  • Full embedding matrix: 519.31 MB

Per-session bandwidth

Workload Prompts Total tokens Unique rows Columnar bytes Dense ship Ratio Saved
qa_short 5 33 27 94.50 KB 519.31 MB 0.0178% 99.9822%
qa_medium 3 68 53 185.50 KB 519.31 MB 0.0349% 99.9651%
code_python 3 81 49 171.50 KB 519.31 MB 0.0323% 99.9677%
chat_turn 1 66 53 185.50 KB 519.31 MB 0.0349% 99.9651%
long_context 1 628 109 381.50 KB 519.31 MB 0.0717% 99.9283%
multilingual 5 38 31 108.50 KB 519.31 MB 0.0204% 99.9796%

Multi-session amortization

Simulating 100 independent sessions on one node with top-100 popular tokens pre-warmed in the row cache. (Realistic for a multi-tenant edge node where BOS/EOS/common BPE pieces saturate quickly.)

Metric Value
Sessions 100
Warm-set size 100 rows
Warm-seed bytes (one-time) 350.00 KB
Cold rows fetched (total) 185
Cold bytes fetched (total) 647.50 KB
Columnar total (warm-seed + cold) 997.50 KB
Dense ship per session × N 50.71 GB
Ratio 0.001876%
Bandwidth saved 99.998124%

Notes

  • Rows are exact (no spectral approximation): columnar fetch is a zero-quality-loss alternative to lossy GKQ rank-K SVD for the embedding tensor.
  • These numbers are for the input embedding half. With tied embeddings (Qwen-style) the same matrix is the lm_head; see the design doc, §5, for the speculative top-k strategy that extends the win to the output projection.
  • Multi-session amortization assumes one node serves all sessions. Sharded across N nodes, each node's warm-set + cold-rows scales with its session subset; the ratio against dense-ship-per-node stays comparable.