training_algorithm parameter: supervised fine-tuning (sft, the default) and the reinforcement-learning methods GRPO (grpo) and DPO (dpo). See the LLM fine-tuning guide for when to use each and how to format datasets.
Training job lifecycle
Every training job moves through a fixed sequence of states:1
requested
Your job has been accepted and is queued for execution. Pioneer is allocating compute.
2
running
Training is actively executing. You can stream logs to monitor progress.
3
complete
Training finished successfully. Metrics (F1, precision, recall) are available on the job record, and checkpoints are ready to download.
POST /felix/training-jobs/:id/stop).
Key parameters
base_model is required. Omitting it returns a 422 validation error. The value must be a model ID from GET /base-models or a checkpoint UUID — not a free-form string.Starting a training job
id — you’ll use it to poll status, retrieve metrics, and run inference against your trained model.
Polling status and reading metrics
Poll the job endpoint until status iscomplete or failed:
Stopping a job
If you need to cancel a running job:stopped. Partial checkpoints saved before the stop may still be available.
Checkpoints and downloading weights
Pioneer saves checkpoints during training. You can list them at any point after the job starts:base_model value in a new training job to continue training from that checkpoint.