Python mastery engine Qwen3.8-27B
Verified micro-improvement loop. Unsloth/QLoRA is the preferred Qwen3.8 stack; DPO learns from executable chosen/rejected pairs. Nothing is promoted unless it beats the golden checkpoint without regression.
βgolden executable score
βchat corrections awaiting review
βtraining stack
| Golden adapter | β |
|---|---|
| Model | Qwen/Qwen3.8-27B |
| Precision | 4-bit QLoRA |
| Protected passes | β |
| Latest self-improvement | β |
| Learning mode | β |
| Team learning | β |
Owner corrections from chat are stored as pending evidence first. Team Mode can review them before they become training pairs.
Training dataset
Built automatically from every teacher-escalated answer + verified skill. It grows as you teach Fleet AI.
βunique examples
βready to train
200recommended min
GPU status β
The shared Vast controller. Fleet only ever wakes it for training β never automatically.
| State | β |
|---|---|
| Reserved by | β |
| Purpose | β |
| Hold remaining | β |
| Rate | β |
General knowledge LoRA workshop
Separate from the Qwen3.8 Python mastery engine above. This legacy/general workshop fine-tunes the current local assistant dataset and remains manual only.
Safety: if a customer production build holds the GPU, this shows busy and refuses β never a takeover. While you're training, customer TTS uses its OpenAI/Kokoro fallback. Idle-stop is paused while your hold is active. Trained adapters always archive to Wasabi (never only on the ephemeral box).
Recent training runs
No runs yet.