Skip to content

Latest commit

 

History

History
197 lines (102 loc) · 8.9 KB

File metadata and controls

197 lines (102 loc) · 8.9 KB

Examples

We have prepared a collection of example notebooks for expressing the toolkit's functionality.

  • algorithms/ contain demonstrations of the toolkit's built-in algorithms. Its generics/ subfolder illustrates config-based generic controls and demonstrates how modular controls can be constructed, and its wrappers/ subfolder covers the wrappers around existing libraries (e.g., trl, mergekit).
  • recipes/ are worked examples that compose existing toolkit components into something new.
  • studies/ demonstrate more extensive studies that compare methods on a given use case.

Algorithms

Algorithm notebooks demonstrate how each method (i.e., control) operates. The methods are grouped below by the category of the model or generation process they act on.

  • Input control


    Input control methods adapt the input (prompt) before the model is called, for example by rewriting or augmenting it or by supplying in-context examples. Current notebooks cover:

    :octicons-arrow-right-24: CPO

    :octicons-arrow-right-24: FewShot

    :octicons-arrow-right-24: GEPA

    :octicons-arrow-right-24: PRewrite

    :octicons-arrow-right-24: SystemPrompt

  • Structural control


    Structural control methods adapt the model's weights or architecture, such as by fine-tuning or merging checkpoints. These notebooks use our wrappers around established training and merging libraries. Current notebooks cover:

    :octicons-arrow-right-24: MergeKit wrapper

    :octicons-arrow-right-24: TRL wrapper

  • State control


    State control methods influence the model's internal states (activations, attention, and similar) at inference time. Current notebooks cover:

    :octicons-arrow-right-24: ActAdd

    :octicons-arrow-right-24: AngularSteering

    :octicons-arrow-right-24: CAA

    :octicons-arrow-right-24: CAST

    :octicons-arrow-right-24: DirectionalAblation

    :octicons-arrow-right-24: ITI

    :octicons-arrow-right-24: PASTA

  • Output control


    Output control methods influence the model's behavior at generation time through the generate() method, by shifting logits, searching over candidates, or shaping the decoding process. Current notebooks cover:

    :octicons-arrow-right-24: BestOfN

    :octicons-arrow-right-24: BudgetForcing

    :octicons-arrow-right-24: ContrastiveDecoding

    :octicons-arrow-right-24: DeAL

    :octicons-arrow-right-24: DExperts

    :octicons-arrow-right-24: RAD

    :octicons-arrow-right-24: SASA

Generic controls

Several of the methods above are specific settings of a smaller number of generic controls. As part of the toolkit, we have prepared a collection of such config-based controls, which we call generics, to enable custom construction of (modular) controls.

The notebooks below show how to configure each generic (as well how to use them to build some of the named controls).

  • State control


    The composable activation-steering atom; each adapter wires a transform, layer selection, and optionally a gate and token scope into one single-behavior control. Current notebooks cover:

    :octicons-arrow-right-24: ActivationAdapter

  • Output control


    The output analogues, one generic per shape: per-candidate value shifts, mixed log-prob sources, segment search, phased splicing, and stop rules. Current notebooks cover:

    :octicons-arrow-right-24: ValueGuidance

    :octicons-arrow-right-24: ContrastiveGuidance

    :octicons-arrow-right-24: SearchDecoding

    :octicons-arrow-right-24: PhasedDecoding

    :octicons-arrow-right-24: StoppingRules

Recipes

Recipes describe useful applications/compositions of the toolkit's functionality. Generally, recipes are where non-trivial combinations of steering methods (beyond the named controls) are demonstrated.

  • Honest-persona prompting


    This notebook reproduces some of the honest-only persona prompting from Anthropic's evaluating honesty post by composing UserPrefix (the |HONEST_ONLY| control token), SystemPrompt (the mode definition), PhasedDecoding (the <honest_only> tag prefill), and StoppingRules (the closing-tag stop). The notebook compares three prompt variants against the (unsteered) baseline on a scenario that pressures the model to misstate a fact.

    :octicons-arrow-right-24: See the recipe

  • Routed decoding


    This notebook fits calibrated probes (ProbeSet) on contrastive prompt pools, combines them with boolean routing rules, and routes each query to a response strategy (a canned response, a disclaimer-prefixed answer, or plain generation) via the RoutedDecoding driver.

    :octicons-arrow-right-24: See the recipe

  • Sharing pipelines (.spipe)


    This notebook fits a CAA control, freezes the steered pipeline into a portable .spipe bundle (the recipe plus the fitted artifacts, content-addressed), and reconstructs the pipeline from the file alone with matching greedy generations.

    :octicons-arrow-right-24: See the recipe

  • Serving through a vLLM server


    This notebook fits a CAA direction in process, saves the SteeringVector, and serves it through a vLLM server running the vLLM-Hook plugin via the vllm-serve backend. The served pipeline holds no model, and its generations are compared against an unsteered pipeline on the same server.

    :octicons-arrow-right-24: See the recipe

Studies

Studies provide in-depth comparisons of steering methods on a given use case. Note that these notebooks can be computationally heavy.

  • :material-list-box-outline: Instruction following


    This notebook studies the effect of post-hoc attention steering (PASTA) on a model's ability to follow instructions, on single-instruction prompts from Split-IFEval. The Inspect task scores each response with the strict IFEval checker and a reward-model quality score, and delivers each prompt's instruction lines to PASTA through per-sample runtime kwargs. We sweep the steering strength and investigate the trade-off between instruction following and response quality.

    :octicons-arrow-right-24: See the study

  • :material-comment-question-outline: Commonsense MCQA


    This notebook studies steering methods on the CommonsenseQA dataset, comparing a few-shot sweep against a DPO-trained LoRA adapter and the unsteered baseline. The Inspect task measures accuracy and positional bias under deterministic choice shuffling; the notebook sweeps the number of few-shot examples and composes the figures from the library plotting calls.

    :octicons-arrow-right-24: See the study

  • :material-call-split: Routing versus prompting


    This notebook compares the probe-based routing from the routed decoding recipe against two prompting baselines that desribe the same referral policy, i.e., the full policy in a system prompt and a prompted classifier.

    :octicons-arrow-right-24: See the study