SGT Fleet

SGT Fleet is a free orchestration runtime that routes defined roles to frontier conductors and specialist workers. Maintained launch OOP and Skills target DeepSeek Flash v4.1 only. JEV procedure retrieval is now integrated.

Software
Free to download and run.
Hosted retrieval
Available on paid plans.
Encrypted local edition
Planned, not shipped yet.
Provider fees
Your Claude, OpenAI and Ollama subscriptions and CLI usage are billed separately by those providers.

Optional structured decision layer

JEV turns task state into a bounded routing decision.

SGT gives JEV the current task state and a typed question. JEV returns a structured answer—such as a route, score or probability—so SGT can choose procedures or a pre-approved workflow without asking the frontier model to reason through every routing step. JEV does not write code, call tools or execute the Fleet.

JEV decidesSGT executesFrontier approves
ChoiceSelects one option from an allowed list and returns probabilities plus confidence.
ScoreEvaluates supplied content against a rubric and returns a bounded score with confidence.
NoulReturns the probability that a true-or-false statement is correct.

1. OOP skill coverage routing

JEV checks whether the deterministic procedure set covers the task, then ranks bounded additions.

JEV assisted procedure routingSGT builds a deterministic baseline. JEV checks coverage and ranks candidates. SGT may append at most two complete procedure groups before the Fleet executes and the frontier model verifies. If JEV is unavailable, SGT keeps the baseline unchanged. Task stateGoal, criteria,available procedures SGT baselineDeterministic skillskept in order JEV checkCoverage score +candidate ranking SGT policyBaseline plus ≤2complete groups Run + verifyFleet executes;frontier reviews Disabled, error or oversize → baseline unchanged JEV assisted procedure routingSGT builds a deterministic baseline. JEV checks coverage and ranks candidates. SGT may append at most two complete procedure groups. The Fleet executes and the frontier model verifies. Failures retain the baseline. Task stateGoal, criteria, available procedures SGT baselineDeterministic skills stay in order JEV checkCoverage score + candidate ranking SGT policyBaseline plus ≤2 complete groups Run + verifyFleet executes; frontier reviews Failure or opt-out → baseline unchanged
Fail-safe behavior: JEV can add bounded context; it cannot erase the deterministic baseline or bypass SGT’s policy.

2. Optional workflow orchestration

When a user or frontier model designates a repeated workflow, the frontier defines 2–16 legal routes including HOLD. JEV selects among them from fresh state.

Optional JEV Mind orchestrationThe frontier model defines legal workflow templates including a hold route. Fresh state and those templates are sent to JEV. JEV returns one selected route or hold without dispatching. The SGT host validates authority and constraints, executes an accepted route, and the frontier model performs final acceptance. Legal routes2–16 frontiertemplates + HOLD Fresh stateFacts, constraints,allowed actions JEV ChoiceOne route or HOLDdispatch: false SGT validatesAuthority, budget,grounding, quality ExecuteAccepted route;frontier accepts JEV proposes a permitted path. It never grants itself permission. Optional JEV Mind orchestrationFrontier-authored legal routes and fresh state go to JEV. JEV returns one route or hold without dispatching. SGT validates it, executes an accepted route, and the frontier model performs final acceptance. Legal routes2–16 frontier templates + HOLD Fresh stateFacts, constraints, allowed actions JEV ChoiceOne route or HOLD · dispatch: false SGT validatesAuthority, budget, grounding, quality Execute + acceptSGT runs; frontier performs review JEV proposes. SGT enforces.
Permission boundary: the JEV receipt is a decision record, not a dispatch command. The host can execute, reject or hold.

Why it matters: common routing decisions can be answered through compact, typed outputs while the expensive frontier model remains focused on authoring plans, handling exceptions and judging final quality. Actual savings depend on the workflow and provider usage; SGT does not publish a measured JEV savings percentage yet.

Data, opt-in and fallback details

JEV is optional and uses the customer’s own TypeSafe API credential under TypeSafe’s separate terms. When enabled, the integration sends the task state, acceptance criteria and bounded candidate text needed for the decision to the TypeSafe API.

If JEV is disabled, unavailable, invalid or the request is too large, SGT keeps its deterministic baseline. The existing $5 Flash scope is unchanged.

Official patterns: intent routing and confidence-gated routing.

TypeSafe AI and JEV are trademarks of their respective owner. This independent DreamSleepAI, Inc. integration does not imply partnership, endorsement or sponsorship.

Measured results

Visual quality, compared with stock sub-agents

Historical fixed-render study, 1 September 2026: five ratings of the same renders

In one historical fixed-render visual study dated 1 September 2026, the trained SGT Fleet arm scored 7.44 / 10 against stock Sonnet at 6.72 / 10 — a gain of +0.72 points in this test. A frontier hand-build reference scored 6.38 / 10 and is shown for context, not as a competing product.

Boundary: five ratings of fixed renders, not five independent builds. Visual quality only.

  • Trained SGT Fleet7.44/10
    +0.72 points vs stock Sonnet in this test
  • Stock Sonnet6.72/10
  • Frontier hand-build (reference)6.38/10
    Shown for context; not a competing product.
Method and full scores — 1 September 2026 study

What was measured

Historical cinematic final five three-way judgments, study dated 1 September 2026. Five ratings of fixed renders — the same renders rated each time, not five independently generated builds. Visual quality only; this is not a runtime-correctness or general-ability measure.

Arms

  • A — Frontier hand-build
  • B — Trained SGT c2
  • C — Stock Sonnet; exact version not established from the retained report

All five raw sample scores

Sample A (frontier) B (SGT c2) C (stock Sonnet)
16.37.46.9
26.77.26.5
36.07.56.5
46.47.46.7
56.57.77.0
Mean6.387.446.72

Method

Judge identifier recorded in the retained harness: claude-fable-5. Exact model versions are unverified and are not relabeled here. The study used randomized labels with context-blind intention, but no full old tool trace survives, so it is not asserted as an audited double-blind study. SGT ranked first in 5 of 5 ratings of these fixed renders.

Fairness limits

  • Five ratings of fixed renders, not five independent scene builds.
  • Three-way randomized labels; no balanced order record in the old study.
  • No raw request/tool trace to establish full context isolation retrospectively.
  • Fable5.1 and Sonnet5 exact versions unverified; not relabeled.
  • Later c2 was trained using preceding judge feedback; this comparison is not held-out generalization.
  • Different task, rubric and builds from JANUS; does not invalidate JANUS numbers or establish GPT6 superiority.

Blender study, 16 September 2026: same scene, same orchestrator

Fable 5.1 orchestrated both builds of the same German village brief. With SGT Fleet running on DeepSeek Flash 4.1, the final frame scored 7.66 / 10. Claude Sonnet 5 working alone, no fleet, scored 4.72 / 10. The live gallery render of the same scene scored 7.78 / 10 in the same sitting.

German village hero frame built by SGT Fleet on Flash 4.1
SGT Fleet on Flash 4.1 · 7.66 / 10 · 5 rounds, 45 min, 0.92 M tokens
German village hero frame built by Claude Sonnet 5 alone, no fleet, camera re-framed for display
Sonnet 5 alone, no fleet · 4.72 / 10 · 8 rounds, 3 h 49 min, 1.90 M tokens · same scene, camera pulled back for display after judging (frame as judged)
  • SGT Fleet on Flash 4.17.66/10
    +2.94 points vs Sonnet 5 alone in this test
  • Sonnet 5 alone, no fleet4.72/10
  • Live gallery render (reference)7.78/10
    Shown for context; the bar the fleet build was measured against.

Boundary: one scene, one hero view per arm, eight blind scores per arm from two judge families (Claude Opus and Codex) with sealed label mappings. Visual quality only. The fleet's skills were trained on Flash 4.1; the same fleet on Sonnet 5 scored 6.22, and GPT-5.6 Luna alone scored 7.62, so this is a comparison of configurations, not a ranking of models.

Method and full scores — 16 September 2026 study

What was measured

One German spring village scene, one hero view per arm, judged in a single sealed sitting. Each arm received n = 8 scores — four per judge family (Claude Opus and Codex gpt-reserve) — of the same final fixed renders. These are repeated evaluations of final artifacts, not eight independent builds, and not eight independent scenes. Visual quality only; this is not a runtime-correctness or general-ability measure.

All ten final arms

Final sealed sitting, 16 September 2026. Mean and sd over n = 8 scores per arm (1–10 scale). Gate is the deterministic artifact gate on the final attempt.
Arm Final attempt Gate Mean sd
Site render, GPT-6 Astra orchestrating the SGT-GPT6 fleet on Flash 4.1 (reference)——7.780.91
Fable 5.1 orchestrating SGT-Fleet on Flash 4.1, COMPONENT protocol (house, site, driver authored separately)5PASS7.660.51
Fable 5.1 with GPT-5.6 Luna alone, no fleet4PASS7.620.88
Fable 5.1 with Flash 4.1 alone, no fleet8PASS7.450.61
Fable 5.1 orchestrating the SGT-GPT6 fleet on Flash 4.18PASS7.200.97
Fable 5.1 orchestrating SGT-Fleet on Flash 4.1 after the Blender training port, monolithic script8PASS7.140.71
Fable 5.1 orchestrating SGT-Fleet on Sonnet 58PASS6.220.92
Fable 5.1 orchestrating SGT-Fleet on Flash 4.1 before the port8PASS5.831.04
Fable 5.1 with Sonnet 5 alone, no fleet8PASS4.721.05
Fable 5.1 orchestrating SGT-Fleet on GPT-5.6 Luna8FAIL1.380.30

Method

Identical brief and delivery contract across arms, run in a common Blender 5.2 headless environment, a deterministic artifact gate, hands-on orchestration by Fable 5.1 with rounds until the blind judges put the arm within 0.5 of the reference or attempt 8, and blind judging with sealed label mappings by two judge families. The orchestrator inspected every frame before any PASS. The reference row is the live gallery render of the same scene, shown for context and not as a competing product.

Fairness limits

  • One scene and one hero view per arm; not a general model ranking.
  • n = 8 scores per arm are four per judge family on the same final fixed renders, not eight independent builds.
  • Selected final attempts only; earlier attempts are not scored here.
  • Orchestration-assisted iteration: the orchestrator chose revisions between rounds, so this is not a measure of unaided model output.
  • Provider-specific knowledge: the fleet's stores and skills were minted from Flash 4.1 failure modes, so results do not transfer as a model comparison.
  • This is a comparison of configurations, not causal proof of training gains.
  • No deterministic gate result was recorded for the 7.78 reference; it is the live gallery render, displayed for context.

Interactive showcase

Built by the Fleet. Verified by the orchestrator.

Explore Archer, a cinematic website created through specialist web and 3D agents, generated art, measured repairs and responsive verification.

Archer cinematic private-jet website shown inside a full-screen cabin window Open interactive showcase

SGT Fleet capability showcase

Archer

A scroll-driven luxury concept with a fly-through cabin window, animated aircraft, interactive seat map and live 3D globe.

  • Fleet-written web and Three.js code
  • Codex-generated visual assets
  • Five-viewport visual verification
Experience Archer

Archer is a fictional concept. Brief adapted from Jason Lee’s “Archer” walkthrough; independent implementation by SGT Fleet.

Pricing

Simple on purpose.

The SGT Fleet software is free. Paid plans add hosted access to the skills and OOP query service. Your Claude, OpenAI and Ollama subscriptions are billed separately by those providers.

STUDENTS + EDUCATORS

Free

For enrolled students and teaching staff.

  • Full Fleet software
  • Both launch CLIs
  • Updates included
Request education access No payment collected. Verification is manual and pending; no access is activated automatically and no free entitlement is granted by this request.

ENTERPRISE

$10 / seat / month

For organizations that need per-seat licensing.

  • Per-seat licence
  • Private-server deployment by separate arrangement
  • Invoicing
Start enterprise $10 per seat / month, invoiced.

Maintained launch OOP and Skills target DeepSeek Flash v4.1 only. Codex and Fable are conductor integrations, and Sonnet 5 or GPT-5.6 Luna knowledge is customer-created, separate and not included.

Notes

Practical notes.

What is actually tuned?+

SGT tunes skills and OOP role packages — the instructions, routing and review structure around frontier orchestration. Provider model weights are unchanged.

How do I extend to a new task or model?+

Use your own frontier usage on unseen work, produce examples and checks, benchmark the specialist behaviour, then fine-tune a new skill or OOP role package for future routed tasks.