llm-behavior-eval Documentation

class llm_behavior_eval.DatasetConfig(_case_sensitive: bool | None = None, _nested_model_default_partial_update: bool | None = None, _env_prefix: str | None = None, _env_file: DotenvType | None = PosixPath('.'), _env_file_encoding: str | None = None, _env_ignore_empty: bool | None = None, _env_nested_delimiter: str | None = None, _env_nested_max_split: int | None = None, _env_parse_none_str: str | None = None, _env_parse_enums: bool | None = None, _cli_prog_name: str | None = None, _cli_parse_args: bool | list[str] | tuple[str, ...] | None = None, _cli_settings_source: CliSettingsSource[Any] | None = None, _cli_parse_none_str: str | None = None, _cli_hide_none_type: bool | None = None, _cli_avoid_json: bool | None = None, _cli_enforce_required: bool | None = None, _cli_use_class_docs_for_groups: bool | None = None, _cli_exit_on_error: bool | None = None, _cli_prefix: str | None = None, _cli_flag_prefix_char: str | None = None, _cli_implicit_flags: bool | None = None, _cli_ignore_unknown_args: bool | None = None, _cli_kebab_case: bool | None = None, _cli_shortcuts: Mapping[str, str | list[str]] | None = None, _secrets_dir: PathType | None = None, *, file_path: str, dataset_id: str = '', dataset_type: DatasetType, preprocess_config: PreprocessConfig = PreprocessConfig(max_length=1024, gt_max_length=256, preprocess_batch_size=128), seed: int | None = 42)

DatasetConfig is a configuration class for defining the settings of a dataset.

Attributes:

file_path: Local path or Hugging Face repository used to load the dataset. dataset_id: Stable logical identity used for evaluator selection and results.

Defaults to file_path for backward compatibility.

dataset_type: The type of the dataset, represented as an enum. preprocess_config: Configuration for preprocessing the dataset. seed: The random seed for reproducibility.

classmethod default_dataset_id(value: object, info: ValidationInfo) → object

Use the loading source as the logical identity for legacy callers.

model_config: ClassVar[SettingsConfigDict] = {'arbitrary_types_allowed': True, 'case_sensitive': False, 'cli_avoid_json': False, 'cli_enforce_required': False, 'cli_exit_on_error': True, 'cli_flag_prefix_char': '-', 'cli_hide_none_type': False, 'cli_ignore_unknown_args': False, 'cli_implicit_flags': False, 'cli_kebab_case': False, 'cli_parse_args': None, 'cli_parse_none_str': None, 'cli_prefix': '', 'cli_prog_name': None, 'cli_shortcuts': None, 'cli_use_class_docs_for_groups': False, 'enable_decoding': True, 'env_file': None, 'env_file_encoding': None, 'env_ignore_empty': False, 'env_nested_delimiter': None, 'env_nested_max_split': None, 'env_parse_enums': None, 'env_parse_none_str': None, 'env_prefix': 'bias_dataset_', 'extra': 'forbid', 'json_file': None, 'json_file_encoding': None, 'nested_model_default_partial_update': False, 'protected_namespaces': ('model_validate', 'model_dump', 'settings_customise_sources'), 'secrets_dir': None, 'toml_file': None, 'validate_assignment': True, 'validate_default': True, 'yaml_config_section': None, 'yaml_file': None, 'yaml_file_encoding': None}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class llm_behavior_eval.DatasetPreset(name: str, dataset_ids: tuple[str, ...])

A CLI preset and the canonical dataset assets required to expand it.

dataset_ids contains every Hugging Face dataset repository needed for the preset. Future catalog versions may add revisions and checksums.

class llm_behavior_eval.DatasetType(*values)
class llm_behavior_eval.EvaluateFactory

Class to create and prepare evaluators.

static create_evaluator(eval_config: EvaluationConfig, dataset_config: DatasetConfig) → BaseEvaluator

Creates an evaluator based on the dataset configuration.

Args:

eval_config: EvaluationConfig object containing evaluation settings. dataset_config: DatasetConfig object containing dataset settings.

Returns:

An instance of a class that inherits from BaseEvaluator.

static get_evaluator_family(dataset_id: str) → Literal['bias', 'censorship', 'hallucination', 'prompt-injection', 'refusal']

Resolve a supported dataset identifier to its evaluator family.

Args:

dataset_id: Dataset identifier or supported benchmark alias.

Returns:

The evaluator family that owns the dataset contract.

class llm_behavior_eval.EvaluationConfig(**data: Any)

Configuration for behavior evaluation.

Args:

max_samples: Optional limit on the number of examples to process. Use None to evaluate the full set. batch_size: Batch size for model inference. Depends on GPU memory (commonly 16-64). If None, will be adjusted for GPU limits. sample: Whether to sample outputs (True) or generate deterministically (False). use_4bit: Whether to load the model in 4-bit mode (using bitsandbytes).

This is only relevant for the model under test.

device_map: Device map for model inference. If None, will be set to “auto”. max_answer_tokens: Number of tokens to generate per answer. Typical range is 32-256.

Use None to apply the evaluator-family default at runtime.

pass_max_answer_tokens: Whether to pass max_answer_tokens to the model. model_path_or_repo_id: HF repo ID or path of the model under test (e.g. “meta-llama/Llama-3.1-8B-Instruct”). model_output_dir: Optional override for the model output directory slug under results_dir. model_token: HuggingFace token for the model under test. judge_batch_size: Batch size for the judge model (free-text tasks only). If None, will be adjusted for GPU limits. max_judge_tokens: Number of tokens to generate with the judge model. Typical range is 16-64.

Use None to apply the evaluator-family default at runtime.

judge_path_or_repo_id: HF repo ID or path of the judge model (e.g. “meta-llama/Llama-3.3-70B-Instruct”). judge_token: HuggingFace token for the judge model. Defaults to the value of model_token if not provided. sample_judge: Whether to sample outputs from the judge model (True) or generate deterministically (False).

Use None to apply the evaluator-family default at runtime.

use_4bit_judge: Whether to load the judge model in 4-bit mode (using bitsandbytes).

This is only relevant for the judge model.

inference_engine: Whether to run inference with vLLM instead of transformers. Overrides model_engine and judge_engine arguments. model_engine: Whether to run model under test inference with vLLM instead of transformers. DO NOT combine with the inference_engine argument. judge_engine: Whether to run judge model inference with vLLM instead of transformers. DO NOT combine with the inference_engine argument. vllm_config: vLLM-specific configuration (optional). Only used when inference_engine or model_engine/judge_engine is set to “vllm”. results_dir: Directory where evaluation output files (CSV/JSON) will be saved. enable_thinking: Whether to enable thinking (if supported by tokenizer/model).

Use None to apply the evaluator-family default at runtime (on for refusal, off otherwise). The judge is never run with thinking enabled.

enable_thinking_arg_name: Enable thinking argument name in tokenizer’s apply_chat_template (e.g. ‘enable_thinking’). thinking_start_token: Thinking start token to use for the model (e.g. ‘<think>’). thinking_end_token: Thinking end token to use for the model (e.g. ‘</think>’). exclude_thinking_trace_for_judge: Whether to exclude thinking trace from judgement. trust_remote_code: Whether to trust remote code when loading models. sampling_config: Sampling configuration for model inference. mlflow_config: MLflow configuration for tracking (optional). evaluator_family: Evaluator family for the current invocation, if known.

model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

classmethod validate_model_output_dir(value: str | None) → str | None

Validate the model output directory to avoid path traversal attacks and illegal characters.

Args:

value: The model output directory.

Returns:

The validated model output directory.

validate_vllm_config_usage()

Ensure vllm_config is only provided when using vLLM.

class llm_behavior_eval.FreeTextBiasEvaluator(eval_config: EvaluationConfig, dataset_config: DatasetConfig)
evaluate() → None

Run the evaluation process.

This is an abstract method that must be implemented by subclasses.

generate() → Sequence[_GenerationRecord]

Generate the answers for the evaluation.

This is an abstract method that must be implemented by subclasses.

class llm_behavior_eval.FreeTextRefusalEvaluator(eval_config: EvaluationConfig, dataset_config: DatasetConfig)
evaluate() → None

Run the evaluation process.

This is an abstract method that must be implemented by subclasses.

generate() → Sequence[_GenerationRecord]

Generate the answers for the evaluation.

This is an abstract method that must be implemented by subclasses.

class llm_behavior_eval.MlflowConfig(*, mlflow_tracking_uri: str | None = None, mlflow_experiment_name: str | None = None, mlflow_run_name: str | None = None, mlflow_run_id: str | None = None, mlflow_artifact_path_subfolder: str | None = None)

Configuration for MLflow tracking (optional).

Keep this separate from the main EvaluationConfig to avoid coupling MLflow-specific settings with core evaluation logic.

Args:

mlflow_tracking_uri: MLflow tracking server URI (optional). mlflow_experiment_name: Experiment name in MLflow (optional). mlflow_run_name: Run name for new runs (optional). mlflow_run_id: Existing run ID to log to instead of creating a new run (optional). mlflow_artifact_path_subfolder: Optional subfolder under the run for artifacts.

If None, artifacts are logged at run root. Use “timestamp” to auto-append a per-run timestamp subfolder.

model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class llm_behavior_eval.PreprocessConfig(_case_sensitive: bool | None = None, _nested_model_default_partial_update: bool | None = None, _env_prefix: str | None = None, _env_file: DotenvType | None = PosixPath('.'), _env_file_encoding: str | None = None, _env_ignore_empty: bool | None = None, _env_nested_delimiter: str | None = None, _env_nested_max_split: int | None = None, _env_parse_none_str: str | None = None, _env_parse_enums: bool | None = None, _cli_prog_name: str | None = None, _cli_parse_args: bool | list[str] | tuple[str, ...] | None = None, _cli_settings_source: CliSettingsSource[Any] | None = None, _cli_parse_none_str: str | None = None, _cli_hide_none_type: bool | None = None, _cli_avoid_json: bool | None = None, _cli_enforce_required: bool | None = None, _cli_use_class_docs_for_groups: bool | None = None, _cli_exit_on_error: bool | None = None, _cli_prefix: str | None = None, _cli_flag_prefix_char: str | None = None, _cli_implicit_flags: bool | None = None, _cli_ignore_unknown_args: bool | None = None, _cli_kebab_case: bool | None = None, _cli_shortcuts: Mapping[str, str | list[str]] | None = None, _secrets_dir: PathType | None = None, *, max_length: int = 1024, gt_max_length: int = 256, preprocess_batch_size: int = 128)

PreprocessConfig is a configuration class for defining the settings of a dataset preprocessing including tokenization, batching and the train labels.

Attributes:

max_length: The maximum length of the text data. gt_max_length: The maximum length for ground truth data. preprocess_batch_size: The batch size for preprocessing the dataset.

model_config: ClassVar[SettingsConfigDict] = {'arbitrary_types_allowed': True, 'case_sensitive': False, 'cli_avoid_json': False, 'cli_enforce_required': False, 'cli_exit_on_error': True, 'cli_flag_prefix_char': '-', 'cli_hide_none_type': False, 'cli_ignore_unknown_args': False, 'cli_implicit_flags': False, 'cli_kebab_case': False, 'cli_parse_args': None, 'cli_parse_none_str': None, 'cli_prefix': '', 'cli_prog_name': None, 'cli_shortcuts': None, 'cli_use_class_docs_for_groups': False, 'enable_decoding': True, 'env_file': None, 'env_file_encoding': None, 'env_ignore_empty': False, 'env_nested_delimiter': None, 'env_nested_max_split': None, 'env_parse_enums': None, 'env_parse_none_str': None, 'env_prefix': 'bias_preprocess_', 'extra': 'forbid', 'json_file': None, 'json_file_encoding': None, 'nested_model_default_partial_update': False, 'protected_namespaces': ('model_validate', 'model_dump', 'settings_customise_sources'), 'secrets_dir': None, 'toml_file': None, 'validate_default': True, 'yaml_config_section': None, 'yaml_file': None, 'yaml_file_encoding': None}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class llm_behavior_eval.SamplingConfig(_case_sensitive: bool | None = None, _nested_model_default_partial_update: bool | None = None, _env_prefix: str | None = None, _env_file: DotenvType | None = PosixPath('.'), _env_file_encoding: str | None = None, _env_ignore_empty: bool | None = None, _env_nested_delimiter: str | None = None, _env_nested_max_split: int | None = None, _env_parse_none_str: str | None = None, _env_parse_enums: bool | None = None, _cli_prog_name: str | None = None, _cli_parse_args: bool | list[str] | tuple[str, ...] | None = None, _cli_settings_source: CliSettingsSource[Any] | None = None, _cli_parse_none_str: str | None = None, _cli_hide_none_type: bool | None = None, _cli_avoid_json: bool | None = None, _cli_enforce_required: bool | None = None, _cli_use_class_docs_for_groups: bool | None = None, _cli_exit_on_error: bool | None = None, _cli_prefix: str | None = None, _cli_flag_prefix_char: str | None = None, _cli_implicit_flags: bool | None = None, _cli_ignore_unknown_args: bool | None = None, _cli_kebab_case: bool | None = None, _cli_shortcuts: Mapping[str, str | list[str]] | None = None, _secrets_dir: PathType | None = None, *, do_sample: bool | None = None, temperature: float | None = None, top_p: float | None = 1.0, top_k: int | None = 0, seed: int | None = 42)

Configuration for sampling.

Args:

do_sample: Whether to sample from the model. False forces greedy decoding. temperature: The temperature to use when sampling is enabled. top_p: The top-p value for sampling. top_k: The top-k value for sampling. seed: The seed for sampling.

model_config: ClassVar[SettingsConfigDict] = {'arbitrary_types_allowed': True, 'case_sensitive': False, 'cli_avoid_json': False, 'cli_enforce_required': False, 'cli_exit_on_error': True, 'cli_flag_prefix_char': '-', 'cli_hide_none_type': False, 'cli_ignore_unknown_args': False, 'cli_implicit_flags': False, 'cli_kebab_case': False, 'cli_parse_args': None, 'cli_parse_none_str': None, 'cli_prefix': '', 'cli_prog_name': None, 'cli_shortcuts': None, 'cli_use_class_docs_for_groups': False, 'enable_decoding': True, 'env_file': None, 'env_file_encoding': None, 'env_ignore_empty': False, 'env_nested_delimiter': None, 'env_nested_max_split': None, 'env_parse_enums': None, 'env_parse_none_str': None, 'env_prefix': 'bias_sampling_', 'extra': 'forbid', 'json_file': None, 'json_file_encoding': None, 'nested_model_default_partial_update': False, 'protected_namespaces': ('model_validate', 'model_dump', 'settings_customise_sources'), 'secrets_dir': None, 'toml_file': None, 'validate_default': True, 'yaml_config_section': None, 'yaml_file': None, 'yaml_file_encoding': None}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class llm_behavior_eval.VllmConfig(*, max_model_len: int | None = None, judge_max_model_len: int | None = None, tokenizer_mode: Literal['auto', 'slow', 'mistral', 'custom'] | None = None, config_format: str | None = None, load_format: str | None = None, gpu_memory_utilization: float = 0.8, enable_lora: bool = False, max_lora_rank: int = 128, language_model_only: bool = False, enforce_eager: bool = True)

Configuration for vLLM inference.

Keep this separate from the main EvaluationConfig to avoid coupling vLLM-specific settings with core evaluation logic. Only used when inference_engine or model_engine/judge_engine is set to “vllm”.

Args:

max_model_len: Maximum model length for vLLM model inference (optional). judge_max_model_len: Maximum model length for vLLM judge inference (optional).

Defaults to the same value as max_model_len if not specified.

tokenizer_mode: Tokenizer mode forwarded to vLLM (e.g. ‘auto’, ‘slow’, ‘mistral’, ‘custom’). config_format: Model config format hint forwarded to vLLM (optional). load_format: Checkpoint load format hint forwarded to vLLM (optional). enable_lora: Whether to enable LoRA. max_lora_rank: The maximum LoRA rank (do not set too high to avoid wasting memory). language_model_only: Whether evaluated-model loads omit multimodal encoders

and load only the language model. Judge loads always omit multimodal encoders. This setting has no effect on text-only architectures.

enforce_eager: Whether to enforce eager execution (useful for CPU-only setups or for saving memory on CUDA graphs).

model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

llm_behavior_eval.expand_dataset_preset(name: str) → list[str]

Expand a supported CLI preset into canonical dataset identifiers.

Args:
name: Case-insensitive supported preset key, such as

bloom:bias:gender:ambiguous or refusal:all.

Returns:

list[str]: Canonical dataset identifiers required by the preset.

Raises:

ValueError: If name is not a supported preset key.

llm_behavior_eval.list_dataset_presets() → tuple[DatasetPreset, ...]

Return every supported preset and its complete asset expansion.

Returns:

tuple[DatasetPreset, …]: Every supported preset and its complete asset expansion in deterministic catalog order.

llm_behavior_eval.load_tokenizer_with_transformers(model_name: str, token: str | None = None, trust_remote_code: bool = False) → TokenizersBackend | SentencePieceBackend

Load a tokenizer by first trying the standard method and, if a ValueError is encountered, retry loading from a local path.

Args:

model_name: The repo-id or local path of the model to load. token: The HuggingFace token to use for the model. trust_remote_code: Whether to trust remote code.

llm_behavior_eval.load_transformers_model_and_tokenizer(model_name: str, token: str | None = None, use_4bit: bool = False, device_map: str | dict[str, int] | None = 'auto', trust_remote_code: bool = False) → tuple[PreTrainedTokenizerBase, GenerativePreTrainedModel]

Load a tokenizer and a causal language model based on the model name/path, using the model’s configuration to determine the correct class to instantiate.

Optionally load the model in 4-bit precision (using bitsandbytes) instead of the default 16-bit precision.

Args:

model_name: The repo-id or local path of the model to load. token: The HuggingFace token to use for the model. use_4bit: If True, load the model in 4-bit mode using bitsandbytes. device_map: The device map to use for the model. trust_remote_code: Whether to trust remote code.

Returns:

A tuple containing the loaded tokenizer and model.

llm_behavior_eval.pick_best_dtype(device: str, prefer_bf16: bool = True) → dtype
Robust dtype checker that adapts to the hardware:
  • chooses bf16→fp16→fp32 automatically

Local and offline datasets

DatasetConfig.file_path is the physical loading source and may be a local directory. Set DatasetConfig.dataset_id to the canonical dataset identifier so evaluator dispatch, prompts, output names, metrics, and provenance match the online preset. If omitted, dataset_id defaults to file_path.

DatasetConfig(
    file_path="/opt/assets/halueval",
    dataset_id="hirundo-io/halueval",
    dataset_type=DatasetType.BIAS,
)

Release tooling can call list_dataset_presets() to enumerate every supported preset and its complete dataset_ids expansion. expand_dataset_preset() provides the same expansion used by the CLI.

Public catalog of supported behavior-evaluation dataset presets.

class llm_behavior_eval.presets.DatasetPreset(name: str, dataset_ids: tuple[str, ...])

A CLI preset and the canonical dataset assets required to expand it.

dataset_ids contains every Hugging Face dataset repository needed for the preset. Future catalog versions may add revisions and checksums.

llm_behavior_eval.presets.build_bias_dataset_id(prefix: str, bias_type: str, kind: str) → str

Build the canonical identifier for a bias dataset asset.

Args:

prefix: Dataset family, such as bbq or bloom. bias_type: Bias category, such as age or gender. kind: Dataset variant, such as bias or unbias.

Returns:

str: Canonical Hugging Face dataset identifier.

llm_behavior_eval.presets.expand_dataset_preset(name: str) → list[str]

Expand a supported CLI preset into canonical dataset identifiers.

Args:
name: Case-insensitive supported preset key, such as

bloom:bias:gender:ambiguous or refusal:all.

Returns:

list[str]: Canonical dataset identifiers required by the preset.

Raises:

ValueError: If name is not a supported preset key.

llm_behavior_eval.presets.list_dataset_presets() → tuple[DatasetPreset, ...]

Return every supported preset and its complete asset expansion.

Returns:

tuple[DatasetPreset, …]: Every supported preset and its complete asset expansion in deterministic catalog order.