Harness notes
Live public comparison from Option Table · Open source table. Scores are research judgments, not measured task outcomes.
{{ catalogError }}
Loading notes…
| Harness | Best fit | Automation interface | Tradeoff |
|---|---|---|---|
| {{ item.name }} | {{ item.best_for }} | {{ item.automation }} | {{ item.tradeoff }} |
Unknown harness note.
{{ item.name }} · implementation notes
{{ item.summary }}
- {{ field }}
- {{ item[field] }}
{{ item.license }} · source revision {{ item.commit }}
{{ capture.provenance }} {{ capture.observation }}
Session history and recovery
{{ catalog.history_semantics.summary }}
- {{ item.harness }} · {{ item.status }}
- {{ item.note }}
Supervising a harness · persistence and lifecycle
{{ catalog.supervisor.summary }}
- {{ item.state }} · {{ item.survives }}
- {{ item.owner }}. {{ item.note }}
{{ item.name }}
{{ item.initial }}
{{ item.resume }}{{ item.persistence }} {{ item.session_id }}
Interface source{{ benchTab==='memory'?'Memory':benchTab==='trials'?'Vendor trials':benchTab==='safety'?'Safety blocks':benchTab==='unavailable'?'Observation unavailable':benchTab==='hindsight'?'Hindsight':'Session diagnostics' }}
Sign in to inspect private FairyStack session evidence.
{{ workspaceStatus }} {{ authError }}
{{ evidenceError }}
{{memoryError}}
{{memoryNotes.stance}}
{{memoryNotes.case_title}}
- {{note.label}}
- {{note.text}}
Design questions
- {{question}}
| Approach | Mechanism | Open question |
|---|---|---|
| {{option.name}} | {{option.mechanism}} | {{option.question}} |
First experiment
{{memoryNotes.experiment}}
{{memoryNotes.benchmark_note}} LongMemEval paper
Import: {{ evidence.sync.status }} · Last success {{ stamp(evidence.sync.last_success_at) }} · {{ evidence.sync.error || '' }}
| Agent / effort | Turns | Median completed turn | Assessed sessions | Cached input |
|---|---|---|---|---|
| {{ r.model }} · {{ r.effort }}{{ r.runtime }} · {{ r.tier }} | {{ r.turns }} | {{ r.minutes }} min{{ r.timed }} with completion boundaries | {{ r.achieved }} / {{ r.assessed }}assessed achieved · uniform recorded settings | {{ r.cache }}{{ r.cacheCalls }} / {{ r.calls }} calls report cache counts |
No comparable recorded turns for this selection yet.
Effort, unattended work and cache cost
Effort controls reasoning per response. Continuing useful work while you are away also requires an unfinished objective and authority to act. A gap between messages does not establish human absence or wasted time.
At Astra’s standard short-context API prices, replacing 100,000 cached input tokens with uncached input adds $0.90; writing them back to cache adds $1.15 instead. These are hypothetical API-price differences, not measured switch penalties or Codex subscription charges. Prices checked 27 September 2026.
GPT-6’s API can preserve the prefix by appending a configuration update while retaining the original request-level effort. This is documented API behavior; use by our installed Codex runtime has not been verified. OpenAI cache documentation.
Times include tools and waiting within a recorded turn, and exclude turns without explicit completion boundaries. Speed tiers stay separate. Tasks, context length and review requirements differ; these observations cannot establish that an effort setting caused a quality or speed change. Unknown identities remain excluded from the comparison. Session assessments below are historical automated judgments, with unassessed work retained.
Observed work, not matched trials. Turn completion does not prove the goal succeeded. Evidence method · Hindsight · Safety blocks
Effort history is still importing. Missing identities cannot support model comparisons.
| Goal / session | Outcome | Active time | Agent |
|---|---|---|---|
| {{ s.session_summary || s.id }}{{ s.session_app || 'Unattributed' }} | {{ s.goal_evaluation_state==='completed' ? (s.goal_outcome || 'Unassessed') : 'Unassessed' }}{{ s.goal_reason }} | {{ Number.isFinite(s.active_work_seconds) && s.active_work_seconds>=0 ? (s.active_work_seconds/60).toFixed(1)+' min' : 'Not recorded' }} | {{ s.runtime || 'Unknown harness' }}{{ identity(s) }} |
No sessions in this context and window.
What would establish cost effectiveness?
| Benefit | Accepted outcomes on the same task and starting commit; retain failed attempts and later regressions. |
|---|---|
| Cost | All attempts’ model charges, elapsed time, and human review/rework minutes. Subscription allocations stay separate. |
| Comparable combination | Harness version × exact model × effort, same tools and permissions; repeated matched tasks grouped by work type. |
| Decision | Compare acceptance rate, cost per accepted task and time per accepted task. Current observations cannot isolate model effects or establish a winner. |
Retained hindsight · {{ evidence.hindsight.length }} observations
Historical automated assessments, not fresh verification or human acceptance. Several observations can refer to the same goal.
No hindsight in this window. Try All retained.
{{ trialError }}
{{ trials.summary }}
{{ trials.unknown_cost_runs }} attempts have unreconciled billing; reported spend is incomplete.
| Agent / task | State | Hidden checks | Time | Cost |
|---|---|---|---|---|
| {{ run.provider }}{{ run.task_title }} | {{ run.state.replaceAll('_',' ') }}{{ run.note }} | {{ run.passed===null?'Not run':run.passed+' / '+run.total }}{{ run.grade }} | {{ run.duration_seconds===null?'—':Math.round(run.duration_seconds)+'s' }} | {{ run.cost_usd===null?'Unknown':'$'+run.cost_usd.toFixed(4) }} |
Method, budget and reproducibility
{{ trials.method }}
{{ trials.budget_policy }}
{{ trials.baseline_verification }}
Native access probes
{{ JSON.stringify(trials.access_receipt,null,2) }}Pinned runtime checksums
{{ JSON.stringify(trials.runtime_receipts,null,2) }}Baseline check receipt
{{ JSON.stringify(trials.baseline_receipt,null,2) }}- {{ task.title }}
- {{ task.specification }}
{{ trials.limitations }}
{{ provider.name }}: {{ provider.access }} Official interface ↗
{{ run.provider }} · {{ run.task_title }} · evidence
{{ run.evidence }}Fixture {{ run.fixture_sha256 }} · runtime {{ run.runtime_version || 'not recorded' }} · model {{ run.model || 'not recorded' }}
“{{ safetyBlocks.signature }}”
| Session / work | Harness | Evidence |
|---|---|---|
| {{ item.session_id.slice(0,4) }} ↗{{ item.task }} | {{ item.harness }}Model: {{ item.model || 'Unknown' }} | {{ item.evidence_label }} ↗ |
- Shared
- {{ safetyBlocks.shared }}
- Unknown
- {{ safetyBlocks.unknown }}
- Next check
- {{ safetyBlocks.next_check }}