Your Own Task
ct run eval assembles one Inspect Task from a task selection, a policy or protocol, Control Tower's scorers and a sandbox. A project that wants to change one of those parts builds the Task itself from the same pieces. Nothing in ct run eval has to know: there are no task args for "my samples", "no opening message", "no task scorer" or "a plain submit tool", because each of those is a line in your own Task. Run that Task through Inspect's own eval, or hand ct run eval a factory for it with --task-factory and keep the task fan-out, the setting workers, the EC2 fleet and the upload.
The pieces
| Piece | Module | What it does |
|---|---|---|
build_policy_solver | control_tower.protocols.policies.policy_solver | Wraps a registered policy in a solver. kickoff_message drops the sample's blank user turns and appends itself when none of the remaining ones is non-empty (DEFAULT_KICKOFF_MESSAGE, "Begin working on the task.", by default; None or "" leaves the messages alone); submit_config is the submit tool the policy gets (None composes it from the sample's main and side tasks). |
wrap_with_task_setup | control_tower.task.setup | Runs the sample's task setup scripts and baseline snapshots before the solver. |
get_linux_scorers, diagnostic_scorers | control_tower.task.scorers, control_tower.task.diagnostics | The main and side task scorer with diagnostics, or the transcript diagnostics on their own. |
DEFAULT_SUBMIT_CONFIG | control_tower.tools.submit | The no-argument submit tool config. |
SandboxConfig | control_tower.sandbox.cli | apply() configures Control Tower's Docker sandbox class from the network flags; build_sandbox() is the task-level spec when an external sandbox replaces the per-sample ones. |
build_eval_dataset | control_tower.ct_runs.ai_eval.dataset | Control Tower's own samples for a TaskSelectionInput, when you want to start from those and change them. |
DEFAULT_MESSAGE_LIMIT | control_tower.defaults | The message cap ct run eval puts on the Task, which stops a policy that never submits. An Inspect Task has no cap of its own. |
The policy builds its own model from its model argument, but ct run eval also puts that name on the Task, so the eval reports the model the policy is running and eval() needs no --model of its own. A Task that leaves model unset falls back to Inspect's resolution, which is a different model or an error.
Example
A policy eval over your own samples, with the dataset's messages left as they are, the plain submit tool, and only the refusal diagnostic scored:
from inspect_ai import Task, eval from control_tower.defaults import DEFAULT_MESSAGE_LIMIT from control_tower.protocols.policies.policy_solver import build_policy_solver from control_tower.protocols.policies.registry import policies from control_tower.task.diagnostics import diagnostic_scorers from control_tower.task.setup import wrap_with_task_setup from control_tower.tools.submit import DEFAULT_SUBMIT_CONFIG policy = policies.get("honest") policy_args = policies.validate_args(policy, {"model": "sonnet"}) task = Task( dataset=my_dataset, solver=wrap_with_task_setup( build_policy_solver( policy.function, policy_args, kickoff_message=None, submit_config=DEFAULT_SUBMIT_CONFIG, ) ), scorer=diagnostic_scorers(["refusal"], timeout_seconds=600), model=policy_args.model.model_name, message_limit=DEFAULT_MESSAGE_LIMIT, ) eval(task)
Leave kickoff_message and submit_config at their defaults and swap diagnostic_scorers for get_linux_scorers and the Task behaves exactly as a ct run eval policy run does; the only thing you brought is the dataset. policies.get also resolves module.path:function, so the policy can be yours too.
A submit_config you pass replaces the one composed from the sample's tasks, so the submit fields those tasks declare are gone and a task scorer reading them back finds nothing. Override it when you have replaced the task scorers too, as the example does. A policy that adds fields of its own, as honest adds confusions, adds them to the one you passed; a policy whose submission schema is its own output contract, as find_questions, ignores it. Keep its name submit: a monitor sees a tool call under the name it was made with, and compose_submit_config enforces that only for the config it builds.
What your samples carry
Every sample carries SampleMetadata (control_tower.task.sample_metadata) in its metadata: an environment block and, when the sample has them, main_task and side_task, in the shape build_eval_dataset writes. A policy's tools, task setup, the composed submit tool and the task scorer all read it, and resolve the environment and tasks it names from the registry, so those must be registered in the running venv. A sample with only an environment block runs a policy against that environment with no task. Each sample sets its own sandbox.
A sample's input is its opening user turn. The default kickoff_message drops an empty one and supplies its own, so ct run eval samples can leave input blank; a caller who passes kickoff_message=None has taken that over and needs a non-empty input on every sample, because a blank user turn is what strict providers like Moonshot and Kimi reject.
Running it with ct run eval
ct run eval --task-factory module:function replaces control_eval, the factory ct run eval calls by default, with yours. The factory receives exactly control_eval's keyword arguments (task_selection, policy, sandbox, run_name and the rest, serialized) and returns the Task to run. Everything around it is unchanged: the task selection still picks the work, each isolated setting still runs its slice in its own venv, --ec2 still fans out one fleet job per task combination and ships the factory's code with --extra-src, and the merged log is still uploaded as a run.
The shortest factory starts from control_eval and changes what it wants:
from typing import Any from inspect_ai import Task, task from control_tower.inspect_tasks import control_eval from control_tower.task.diagnostics import diagnostic_scorers @task def my_task(**kwargs: Any) -> Task: base = control_eval(**kwargs) return Task( dataset=base.dataset, solver=base.solver, scorer=diagnostic_scorers(["refusal"], timeout_seconds=600), model=base.model, model_roles=base.model_roles, sandbox=base.sandbox, name=base.name, message_limit=base.message_limit, metadata=base.metadata, )
uv run ct run eval --task-factory my_project.tasks:my_task --extra-src ../my-project \ --policy honest --task-file simple --ec2
A factory that composes the pieces itself, as in the first example on this page, reads its selections out of the same keyword arguments: kwargs["task_selection"] validates as TaskSelectionInput and feeds build_eval_dataset, kwargs["policy"] as FunctionSelectionInput names the policy. The factory is decorated with Inspect's @task, and ct run eval refuses one that is not: the local and setting-worker paths run through eval_set, which reuses a log in --log-dir whose task name, recorded arguments and model match, and only a decorated factory records arguments of its own, so an undecorated one would make every run with the same run name look like a rerun of the first, or of the control_eval run it wraps. The run's command metadata records the --task-factory flag. The factory's module has to import in the venv that launches the run, control-tower's plus --extra-src, on --ec2 too: the launcher imports and checks the factory before the fleet provisions, so a factory that only imports inside an isolated setting's venv is rejected there.
What the factory does not change is where the samples come from: they are the task selection's, resolved from registered environments, so the fleet can fan them out and the workers can rebuild them. A dataset with no registered environment behind it is an ordinary Inspect eval, run with eval() as in the first example, or submitted to the fleet as its own job through FleetController.run with the factory as task_ref.