Scaffolding Summary

Summary of 20+ scaffolding campaigns across reasoning (R1-R6), creativity (C1-C5), and forecasting (F1-F4), totaling ~2,500 experiments. Converges on the distinction between.

heuristic-algebra scaffolding-experiments skeletons-vs-programs reasoning creativity forecasting
source Updated 2026-04-12

Scaffolding Summary: From Reasoning to Creativity to Forecasting

Summary of 20+ scaffolding campaigns across reasoning (R1-R6), creativity (C1-C5), and forecasting (F1-F4), totaling ~2,500 experiments. Converges on the distinction between skeleton scaffolds (structure without content) and program scaffolds (structure AND content).

Core Discovery: Skeletons vs Programs

Skeletons (Structure Without Content)

  • Name qualities to exhibit but provide no process instructions or domain content
  • Example: CT scaffold (Clarity, Accuracy, Precision, Relevance, Logical Coherence, Evidential Sufficiency, Intellectual Humility, Fairness)
  • Compose with heuristics: Heuristics fill the skeleton with domain-specific content
  • CT + Economics: +620 Elo; CT + Forecasting: +431 Elo

Programs (Structure AND Content)

  • Provide step-by-step directives with domain-specific instructions
  • Examples: GI (creativity axioms), FS (forecasting principles)
  • Do NOT compose with heuristics: Additional heuristics compete with existing instructions, degrading performance
  • GI + any heuristic: -99 to -196 Elo; FS + any heuristic: -97 to -208 Elo

Reasoning Evolution: Skeleton to Program

  • R1: CT skeleton alone (~1538 Elo)
  • R2: CT ⊕ ProbStats = Calibrated Reasoning (1759 Elo, superadditive)
  • R4-R6: Compounds like Decisive Calibration (1830 Elo), Reflexive Reasoning (1811 Elo)
  • Progression shows skeleton gaining content through composition until becoming self-contained programs

Three-Task Architecture

Task Scaffold Type Composes? Optimal Config
Reasoning CT Skeleton Yes CT + best heuristic
Creativity GI Program No GI alone
Forecasting FS Program No FS alone

Key Findings

  • Programs stronger than enhanced skeletons: GI alone beats CT + any heuristic
  • Rubric must match task: Default rubric rewards structure, creativity rubric rewards novelty
  • Data in prompt shifts rankings: Data-analysis scaffolds (FS) win with actual time series
  • Domain heuristics often fail on their own domain due to redundancy

Scaffold Design Principles

  1. Skeletons compose, programs don’t
  2. Programs are stronger than skeletons
  3. Rubric is part of the experiment
  4. Prompt structure matters (data vs no data)
  5. Domain overlap creates redundancy, not reinforcement

Open Questions

  • Do reasoning compounds (DC, RR) compose with heuristics?
  • Can skeletons be built for creativity/forecasting?
  • Multi-query validation stability
  • Universal skeleton possibility
  • Human evaluation correlation