Context Compression & SQLite Token Telemetry
prompt cost reduction
Published: 2026-06-20 | Project: Bet Bodhi | Discipline: Cognitive AI & Multi-Agent Swarms
Author: Nicholas Alexander MacAskill — Founder & CTO, Flocano Labs | Canonical: https://www.nicholasmacaskill.com/dossier/llm-finops-optimization
Context Compression Algorithm
To optimize LLM prompt structures and mitigate high token usage costs, the compressContext engine parses incoming scraping inputs and retains only blocks carrying specific domain keywords. It then slices the output context to a maximum character size:
export function compressContext(text: string, query: string = "", maxChars: number = 4000): string {
if (!text) return "";
const defaultKeywords = ["drawdown", "rules", "limits", "payouts", "error", "stats", "kelly", "stake", "edge"];
const queryKeywords = query ? query.toLowerCase().split(/[^a-z0-9]+/gi).filter(w => w.length > 3) : [];
const keywords = Array.from(new Set([...defaultKeywords, ...queryKeywords]));
const chunks = text.split(/(?:\r?\n|\. |\<[^\>]+\>)/g).map(c => c.trim()).filter(c => c.length > 0);
const filtered = chunks.filter(chunk => {
const lower = chunk.toLowerCase();
return keywords.some(kw => lower.includes(kw));
});
let compressed = filtered.join("\n") || text;
return compressed.length > maxChars ? compressed.slice(0, maxChars) + "\n... [TRUNCATED] ..." : compressed;
}Budget Circuit Breakers
The TokenTracker intercepts response usage metadata, estimates real-time API fees based on the model's rate structures (e.g. 0.075/0.30 per million tokens for Gemini 2.0 Flash), and records logs to SQLite. If daily spending hits 80% (1.60) or 100% (2.00) of the budget threshold, the tracker triggers an asynchronous warning message via the Telegram Bot API.