SFT Loops & Feedback Moats
Continuous Improvement of LLM Scopes
Published: 2026-06-20 | Project: BayesianPivot | Discipline: Cognitive AI & Multi-Agent Swarms
Author: Nicholas Alexander MacAskill — Founder & CTO, Flocano Labs | Canonical: https://www.nicholasmacaskill.com/dossier/bp-retraining
The Recursive Loop (Outcome): The Feedback Moat
By structuring the database around a unified signed_ledger (for bot signals) and journal (for manual discretionary setups), the system achieves a self-improving feedback loop:
1. Soft Retraining (Few-Shot Context)
Every scan cycle, the system extracts the last 10–20 trades (both successes and failures) from the database and feeds them as contextual examples into the live LLM prompt. The validator instantly learns what patterns are failing in the current market environment and adjusts its scoring threshold (e.g., automatically penalizing similar setups).
2. Hard Retraining (Supervised Fine-Tuning / SFT)
Weekly, the retraining loop outputs instruction-tuned datasets (training_[timestamp].jsonl). These files model the exact conditions of successful "Human Alpha" entries, ready for supervised fine-tuning.
Performance Gains
During optimization runs, SFT integration demonstrated a significant performance lift:
| Retraining Cycle | Win Rate | Average PnL per Trade |
|---|---|---|
| Run 3 (87 samples) | 28.7% | +$4.21 |
| Run 4 (101 samples) | 38.6% | +$57.53 |
Manual discretionary trades labeled ALPHA in the ledger showed a dominant win rate when fading liquidity sweeps. That pattern became the math-only Turtle Soup scanner, shifting the system from reactive indicators to proactive liquidity captures.