We have prepared a collection of example notebooks for expressing the toolkit's functionality.
algorithms/contain demonstrations of the toolkit's built-in algorithms. Itsgenerics/subfolder illustrates config-based generic controls and demonstrates how modular controls can be constructed, and itswrappers/subfolder covers the wrappers around existing libraries (e.g.,trl,mergekit).recipes/are worked examples that compose existing toolkit components into something new.studies/demonstrate more extensive studies that compare methods on a given use case.
Algorithm notebooks demonstrate how each method (i.e., control) operates. The methods are grouped below by the category of the model or generation process they act on.
-
Input control
Input control methods adapt the input (prompt) before the model is called, for example by rewriting or augmenting it or by supplying in-context examples. Current notebooks cover:
:octicons-arrow-right-24: CPO
:octicons-arrow-right-24: FewShot
:octicons-arrow-right-24: GEPA
:octicons-arrow-right-24: PRewrite
:octicons-arrow-right-24: SystemPrompt
-
Structural control
Structural control methods adapt the model's weights or architecture, such as by fine-tuning or merging checkpoints. These notebooks use our wrappers around established training and merging libraries. Current notebooks cover:
:octicons-arrow-right-24: MergeKit wrapper
:octicons-arrow-right-24: TRL wrapper
-
State control
State control methods influence the model's internal states (activations, attention, and similar) at inference time. Current notebooks cover:
:octicons-arrow-right-24: ActAdd
:octicons-arrow-right-24: AngularSteering
:octicons-arrow-right-24: CAA
:octicons-arrow-right-24: CAST
:octicons-arrow-right-24: DirectionalAblation
:octicons-arrow-right-24: ITI
:octicons-arrow-right-24: PASTA
-
Output control
Output control methods influence the model's behavior at generation time through the
generate()method, by shifting logits, searching over candidates, or shaping the decoding process. Current notebooks cover::octicons-arrow-right-24: BestOfN
:octicons-arrow-right-24: BudgetForcing
:octicons-arrow-right-24: ContrastiveDecoding
:octicons-arrow-right-24: DeAL
:octicons-arrow-right-24: DExperts
:octicons-arrow-right-24: RAD
:octicons-arrow-right-24: SASA
Several of the methods above are specific settings of a smaller number of generic controls. As part
of the toolkit, we have prepared a collection of such config-based controls, which we call generics,
to enable custom construction of (modular) controls.
The notebooks below show how to configure each generic (as well how to use them to build some of the named controls).
-
State control
The composable activation-steering atom; each adapter wires a transform, layer selection, and optionally a gate and token scope into one single-behavior control. Current notebooks cover:
:octicons-arrow-right-24: ActivationAdapter
-
Output control
The output analogues, one generic per shape: per-candidate value shifts, mixed log-prob sources, segment search, phased splicing, and stop rules. Current notebooks cover:
:octicons-arrow-right-24: ValueGuidance
:octicons-arrow-right-24: ContrastiveGuidance
:octicons-arrow-right-24: SearchDecoding
:octicons-arrow-right-24: PhasedDecoding
:octicons-arrow-right-24: StoppingRules
Recipes describe useful applications/compositions of the toolkit's functionality. Generally, recipes are where non-trivial combinations of steering methods (beyond the named controls) are demonstrated.
-
Honest-persona prompting
This notebook reproduces some of the honest-only persona prompting from Anthropic's evaluating honesty post by composing
UserPrefix(the|HONEST_ONLY|control token),SystemPrompt(the mode definition),PhasedDecoding(the<honest_only>tag prefill), andStoppingRules(the closing-tag stop). The notebook compares three prompt variants against the (unsteered) baseline on a scenario that pressures the model to misstate a fact. -
Routed decoding
This notebook fits calibrated probes (
ProbeSet) on contrastive prompt pools, combines them with boolean routing rules, and routes each query to a response strategy (a canned response, a disclaimer-prefixed answer, or plain generation) via theRoutedDecodingdriver. -
Sharing pipelines (
.spipe)
This notebook fits a CAA control, freezes the steered pipeline into a portable
.spipebundle (the recipe plus the fitted artifacts, content-addressed), and reconstructs the pipeline from the file alone with matching greedy generations. -
Serving through a vLLM server
This notebook fits a CAA direction in process, saves the
SteeringVector, and serves it through a vLLM server running the vLLM-Hook plugin via thevllm-servebackend. The served pipeline holds no model, and its generations are compared against an unsteered pipeline on the same server.
Studies provide in-depth comparisons of steering methods on a given use case. Note that these notebooks can be computationally heavy.
-
:material-list-box-outline: Instruction following
This notebook studies the effect of post-hoc attention steering (PASTA) on a model's ability to follow instructions, on single-instruction prompts from Split-IFEval. The Inspect task scores each response with the strict IFEval checker and a reward-model quality score, and delivers each prompt's instruction lines to PASTA through per-sample runtime kwargs. We sweep the steering strength and investigate the trade-off between instruction following and response quality.
-
:material-comment-question-outline: Commonsense MCQA
This notebook studies steering methods on the CommonsenseQA dataset, comparing a few-shot sweep against a DPO-trained LoRA adapter and the unsteered baseline. The Inspect task measures accuracy and positional bias under deterministic choice shuffling; the notebook sweeps the number of few-shot examples and composes the figures from the library plotting calls.
-
:material-call-split: Routing versus prompting
This notebook compares the probe-based routing from the routed decoding recipe against two prompting baselines that desribe the same referral policy, i.e., the full policy in a system prompt and a prompted classifier.