Methodology

How Sediment turns coding-agent activity into training data

Sediment turns coding-agent activity into training data by recording model responses, developer decisions, edits, commits, and continuous integration (CI) results. It links these events and selects examples using defined rules.

You can export those examples to train on selected responses, compare preferred and rejected responses, or learn from verified outcomes. Each export records the evidence and rules used to select it.

Capture five signals

Your workflow produces evidence in several systems. Sediment stores each observation as a fact: a record that it appends without changing earlier records. It interprets the evidence later.

  1. Inference call

    A model gateway reports successful model calls, including each request and response. Sediment preserves message order and content types, and links the call to the agent session that made it.

  2. Developer decision

    An agent integration records whether an edit was accepted or rejected. It also distinguishes your explicit decision from an implicit action, such as an edit allowed by configuration.

    Claude Code, GitHub Copilot, Codex, and pi expose different events. Their decision counts aren't directly comparable. For example, pi reports successful edits as implicit accepts, without a human approval signal.

  3. Edit observation

    A hook that runs when a session ends sends the applied model text and the file's state at that point. Sediment later compares them to estimate how much of the edit remains.

    DECODE tracked edits after acceptance that reveal intent missing from acceptance signals and final Git snapshots. Small models fine-tuned on these edits outperformed frontier baselines on the study's edit-prediction tasks.

  4. Push

    A webhook from your Git hosting service reports which commits reached a branch. A local repository mirror stores the code and Git notes that link commits to agent sessions.

  5. CI outcome

    A webhook or CI API reports each run attempt's original result and a standardized result. Capture preserves that evidence without parsing logs or guessing whether a failure was intermittent.

Capture depends on the integrations you enable. Gateway capture supplies model requests and responses. Copilot contributes decision events, but its model calls can't route through the gateway to produce recorded interaction sequences.

How capture works explains each signal and its coverage. The architecture guide describes the network boundaries of a self-hosted deployment.

Turn observations into training examples

Traces, attribution, evaluation, and training-data derivation answer different questions. Keeping them separate lets you inspect both the original evidence and the decisions that produced a training row.

  1. Traces

    Traces record requests, responses, tool calls, and other events. Sediment stores captured observations as facts. A trace alone doesn't establish which response you should train a model to produce.

  2. Attribution

    Attribution links a model call to a commit. Git notes record the contributing sessions; text similarity ranks calls within them. Without a valid note, Sediment estimates the link from token overlap. Each link records the method used.

  3. Evaluation

    Evaluation interprets decisions, edit retention, and CI attempts under defined rules. Sediment records labels and confidence scores so you can assess why an example qualifies for training.

    These labels describe observed outcomes. They don't replace an independent evaluation of the trained model or prove that a response is correct.

  4. Training-data derivation

    Derivation combines the evidence into attributed completions, which collect evidence around one model response, and rollouts, which record a session's sequence of interactions.

    Exports convert these records into training rows. The same facts, repository state, and rules reproduce the result. Changing the rules lets you rebuild the dataset from the original evidence.

Attribution explains Git notes and the similarity fallback. How derivation works explains how Sediment builds the records used by exports.

Choose a training format

Choose a format based on what you want the model to learn. An evidence recipe is a versioned set of rules for selecting examples. Every row names the recipe and the evidence that made it eligible.

A resolved CI result summarizes the recorded attempts for a commit. Here, a clean result has no conflicting workflow verdicts or retry evidence that suggests an intermittent failure.

  • Supervised fine-tuning (SFT)

    SFT teaches a model to produce a selected response. The default curated recipe requires an explicit accept or an accepted edit with a retention score of at least 0.8.

    An explicit reject, abandonment, or a resolved workflow failure excludes the example. A CI pass alone doesn't qualify it. The optional verified recipe requires a clean resolved CI pass. Both recipes require a minimum confidence score.

    Diff-shaped SFT (diff-SFT) applies the same selection rules and uses the exact committed patch as the training target.

    SWE-Gym raised Qwen2.5-Coder-32B-Instruct's SWE-bench Verified resolution rate from 7.0% to 20.6% with OpenHands after training on 491 recorded agent runs that passed task tests: a gain of 13.6 percentage points.

  • Direct preference optimization (DPO)

    DPO teaches a model to favor one response over another. Each pair has a chosen and a rejected response from the same model, with identical prompt histories.

    The default recipe uses explicit accepts and rejects. The optional outcome recipe uses clean CI passes and failures. A pair uses one evidence source. Independent accepts and rejects don't imply that you directly compared the two responses.

  • Reinforcement learning from verifiable rewards (RLVR)

    RLVR uses checkable outcomes as training signals. Sediment exports task descriptions and trajectories: the recorded sequence of agent interactions.

    A resolved CI pass or failure supplies reward evidence. A timeout, skipped run, or unknown result supplies none. Exports support Sediment's audit format and mappings for SWE-bench tasks and NeMo Gym rollouts.

    Sediment exports recorded evidence; it doesn't rerun checks or construct a runnable environment. External trainers can require an adapter and additional task or environment setup.

Inspect an example row

This synthetic SFT example illustrates an explicitly accepted response. It validates against the sft-sample v1 schema. The row is indented for readability; exports use one JSON object per line.

{
  "prompt": [
    {
      "role": "user",
      "content": "Make retry_delay(attempt) return 2 raised to attempt."
    }
  ],
  "completion": [
    {
      "role": "assistant",
      "content": "def retry_delay(attempt):\n    return 2 ** attempt\n"
    }
  ],
  "tools": [],
  "metadata": {
    "org_id": "example-org",
    "source_model": "example-model",
    "completion_id": "example-call-1",
    "recipe_id": "sft_curated",
    "recipe_version": 1,
    "eligibility_source": "explicit_accept",
    "label_confidence": 0.9,
    "ci_reliability": null,
    "provenance": {
      "policy_version": "3",
      "quarantine_revision": 0,
      "policy_digest": null
    },
    "split": "train",
    "schema_id": "https://sediment.so/schemas/training-rows/sft-sample/v1.json",
    "schema_version": 1
  }
}

The trainer uses prompt as context and completion as the response to learn. The tools list is empty because inference-call schema v1 doesn't capture tool definitions.

metadata records the source model, selection rule, confidence, CI reliability, policy, and training or evaluation assignment. Keep this evidence outside the model's training inputs.

Here, ci_reliability: null means the example has no resolved CI evidence. The confidence score expresses the policy's assessment of the label; it isn't a measured probability of correctness.

Use schema_id and schema_version to identify the row format. The repository includes the versioned JSON schemas for validation.

Limitations

Review attribution, redaction, retention, and CI evidence before you use an export for training. Each has limits that affect how you should interpret the data.

  • Attribution uncertainty

    A Git note records a session's contribution to a commit. It doesn't establish exact authorship for every line. Sediment still uses similarity to select model calls within that session.

    Missing hooks, cherry-picks, and squash merges on a hosting service can leave commits without notes. The fallback, called Jaccard similarity, estimates the relationship and reduces confidence according to its similarity score.

  • Redaction

    Before storing content in the fact database, Sediment replaces recognized API keys and bearer credentials with a fixed marker. This uses a fixed pattern set with no operator configuration. It can't detect every secret.

    Redaction at ingestion doesn't remove secrets from existing facts or the repository mirror. Quarantine excludes selected facts from later derivations without rewriting them.

  • Edit survival

    A low retention score doesn't identify who changed the file. The agent's own revision can lower an earlier edit's score. Counts of changes outside the agent's edit tools can also include formatters, Git operations, and shell commands.

    Transcript retention covers Edit and Write calls. External line counts cover Claude Code only. Missing session-end hooks and oversized observations leave gaps. Session-end retention doesn't prove that code later reached a commit.

    SWE-chat found that 44.3% of agent-produced code reached commits in its public, opt-in sample. This includes agent self-overwrites; excluding them, the reported survival rate was 50.3%.

  • CI as a reward

    A CI pass means the recorded checks found no failure. It doesn't prove correctness, and a failure doesn't establish that the code caused it. Timeouts and other results without a pass or failure create no reward.

    Retries that change from failure to pass, or pass to failure, receive zero reliability by default. RLVR can still carry the resulting reward, so check reliability before training. Conflicting workflow verdicts produce no aggregate reward.

    The built-in training and evaluation split keeps each session together. It doesn't separate related tasks, repositories, or identical prompts across sessions. Before training, check that your dataset meets your benchmark's holdout rules.

Sediment leaves unsupported values absent and reports why it skips inputs. Review the export diagnostics to distinguish missing evidence from examples excluded by your rules.

Research sources

These studies support collecting context, corrections, and outcomes. They don't evaluate Sediment or establish gains from its exports. Test whether your dataset improves your model on held-out tasks.

  • DECODE: 53,600 edits to accepted completions from more than 1,000 developers.
  • SWE-Gym: Training agents on recorded runs that passed task tests.
  • SWE-chat: Code retention in 6,000 public, opt-in coding-agent sessions.
  • SWE-ContextBench: Correctly selected summaries improved resolution and reduced tokens and cost; poor selection could hurt accuracy. This studied context reuse without updating model weights.

Next steps

  • Quickstart

    Capture your first agent session on one machine in about five minutes.

  • Export docs

    Choose a training objective, prepare its evidence, and audit the output.

  • How it works

    Read how captured facts become attributed completions and rollouts.

  • Schemas

    Find the fields and validation rules for each export format.