Tagged articles

LLM compression

4 articles · Page 1 of 1
Architecture Digest
Architecture Digest
Sep 13, 2026 · Artificial Intelligence

Running 2.78T-Parameter Kimi K3 on 8GB RAM: Pure C Inference Engine Explained

A pure C inference engine runs the 2.78 trillion parameter Kimi K3 model on just 8GB RAM by streaming weights from disk, leveraging MoE sparsity (only 3.7% active per token) and computing directly on compressed formats, achieving correct output at 32.69 seconds per token on consumer hardware.

BenchmarkingC languageDisk Streaming
0 likes · 11 min read
Running 2.78T-Parameter Kimi K3 on 8GB RAM: Pure C Inference Engine Explained
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Sep 3, 2026 · Artificial Intelligence

LLM Compression of Financial Texts Alters Decisions Despite Factual Accuracy

A new arXiv paper introduces Information Fidelity to measure how LLM compression of 10-Q MD&A sections and earnings calls changes downstream investment decisions, identifying decontextualization and model dependency as key failure modes and proposing Agentic Context Compression (ACC) to audit and reduce decision flips.

10-Q MD&ALLM compressionagentic context compression
0 likes · 22 min read
LLM Compression of Financial Texts Alters Decisions Despite Factual Accuracy
AI Code to Success
AI Code to Success
Mar 27, 2026 · Artificial Intelligence

How Google’s TurboQuant Cuts LLM Memory by 6× and Speeds Up Inference 8×

Google Research’s TurboQuant algorithm compresses large‑language‑model KV caches from 32‑bit to 3‑bit, achieving a six‑fold reduction in memory usage and an eight‑fold inference speedup on H100 GPUs while preserving 100 % accuracy, and it also improves vector search performance without requiring large codebooks.

AI EfficiencyInference AccelerationLLM compression
0 likes · 10 min read
How Google’s TurboQuant Cuts LLM Memory by 6× and Speeds Up Inference 8×