Skip to content

msflib.seed

Part of the msflib core package.

msflib.seed

is_production_db(db_url: str) -> bool

Whether a database URL must be treated as sensitive (conservative heuristic).

Only a SQLite file under ./tmp/ or /tmp/ is local. The path is normalised before it is compared, so sqlite:///./tmp/../prod.db (which SQLite resolves to ./prod.db) is not mistaken for a temp file, a path containing a backslash is never local, and symlinks are followed: a link inside the temp directory that points elsewhere makes the file sensitive. Everything else, including in-memory SQLite and every non-SQLite URL (also PostgreSQL on localhost), is sensitive.

run_seed_cli(*, argv: Sequence[str] | None = None, description: str = 'Run database seeders.', default_seeders_yml: str | None = None, default_data_paths: list[str] | None = None, core_settings: CoreSettings | None = None, import_models: Callable[[], None] | None = None, metadata=SQLModel.metadata, is_production_db: Callable[[str], bool] = is_production_db) -> None

Run database seeders from a YAML configuration file.

This is the main entry point for seeding. It: 1. Parses command-line arguments 2. Confirms sensitive database operations 3. Optionally initializes the database schema 4. Runs seeders from the YAML configuration

Parameters:

Name Type Description Default
argv Sequence[str] | None

Command-line arguments (defaults to sys.argv).

None
description str

Description for the CLI help text.

'Run database seeders.'
default_seeders_yml str | None

Default path to seeders.yml file.

None
default_data_paths list[str] | None

Default list of data directory paths.

None
core_settings CoreSettings | None

Optional settings object that provides database URL if not provided via CLI.

None
import_models Callable[[], None] | None

Optional callable to import all SQLModel models before initialization.

None
metadata

SQLAlchemy MetaData instance (typically SQLModel.metadata).

metadata
is_production_db Callable[[str], bool]

Function to determine if a database URL is likely in production.

is_production_db

base

SeederBase(config: ConfigSchema)

Bases: ABC

Abstract base class for creating seeders.

This class defines the required structure for any seeder implementation. Subclasses must implement both the __init__ and seed methods to handle configuration and seeding logic respectively.

Initialize the seeder with a configuration schema.

Parameters:

Name Type Description Default
config ConfigSchema

Configuration object containing parameters needed for seeding.

required

Example: >>> # You should initialize your config and records as below: >>> self.config = config >>> self.records = ( ... config.records if config.records else [model_action.random().model_dump()] ... )

seed(session: Session, seeds: dict[str, Any]) abstractmethod

Perform seeding logic using the provided database session and seed data.

Parameters:

Name Type Description Default
session Session

SQLAlchemy session object to interact with the database.

required
seeds Dict[str, Any]

Dictionary of seed data to insert into the database.

required

Raises:

Type Description
NotImplementedError

If the method is not implemented by a subclass.

Example: >>> # Always update seeds with the seeded instance. for example: >>> accounts = seeds["accounts"] # getting account instances from the seeds dict >>> profiles = [] >>> for account, record in zip(accounts, self.records): profile = pa.create( session, data=pa.random(**record), update={"id": account.id}, ) profiles.append(profile) >>> # Storing the profile instances so it can be used as a dependency just like >>> # how account instances was used to seed profile. >>> seeds.update({self.config.name: profiles})

ActionSeeder(model_action: ModelActionType, config: ConfigSchema)

Bases: SeederBase

Base class for seeding database tables. Handles: - Fetching dependencies from previously seeded data - Reading data from CSV files - Creating records in the database

Initialize the seeder with the required configurations. Args: model_action (ModelActionType): The action class responsible for creating models. config (ConfigSchema): Configuration containing seeder details.

seed(session: Session, seeds: dict[str, Any])

Seed the database with records. Args: session (Session): The database session for committing transactions. seeds (Dict): This stores the created seeds incase of dependency between seeds.

cli

is_production_db(db_url: str) -> bool

Whether a database URL must be treated as sensitive (conservative heuristic).

Only a SQLite file under ./tmp/ or /tmp/ is local. The path is normalised before it is compared, so sqlite:///./tmp/../prod.db (which SQLite resolves to ./prod.db) is not mistaken for a temp file, a path containing a backslash is never local, and symlinks are followed: a link inside the temp directory that points elsewhere makes the file sensitive. Everything else, including in-memory SQLite and every non-SQLite URL (also PostgreSQL on localhost), is sensitive.

run_seed_cli(*, argv: Sequence[str] | None = None, description: str = 'Run database seeders.', default_seeders_yml: str | None = None, default_data_paths: list[str] | None = None, core_settings: CoreSettings | None = None, import_models: Callable[[], None] | None = None, metadata=SQLModel.metadata, is_production_db: Callable[[str], bool] = is_production_db) -> None

Run database seeders from a YAML configuration file.

This is the main entry point for seeding. It: 1. Parses command-line arguments 2. Confirms sensitive database operations 3. Optionally initializes the database schema 4. Runs seeders from the YAML configuration

Parameters:

Name Type Description Default
argv Sequence[str] | None

Command-line arguments (defaults to sys.argv).

None
description str

Description for the CLI help text.

'Run database seeders.'
default_seeders_yml str | None

Default path to seeders.yml file.

None
default_data_paths list[str] | None

Default list of data directory paths.

None
core_settings CoreSettings | None

Optional settings object that provides database URL if not provided via CLI.

None
import_models Callable[[], None] | None

Optional callable to import all SQLModel models before initialization.

None
metadata

SQLAlchemy MetaData instance (typically SQLModel.metadata).

metadata
is_production_db Callable[[str], bool]

Function to determine if a database URL is likely in production.

is_production_db

runner

SeedContext(seeds: dict[str, list[ModelBase]] = dict(), skipped_seeders: list[str] = list(), seed_stats: Counter = Counter()) dataclass

Stores seeded entities and cumulative run metadata for repeated seed runs.

SeedRunner(seeders_yml_path: str, data_path: str | None = None, data_paths: list[str] | None = None, models_to_seed: list[str] | None = None, show_log: bool = True)

Handles database seeding by running seeders defined in a YAML configuration file.

The SeedRunner loads seeder configurations, processes record data, dynamically imports custom seeder classes, actions, and models, and executes seeding operations into the database. It also allows selective execution of seeders for specific models.

Attributes:

Name Type Description
context SeedContext

Shared context storing seeds and cumulative stats.

seeders dict

Seeder configurations parsed from the seeders YAML file.

data_paths list[str]

Directory paths where seeder data files are searched.

models_to_seed list[str]

Optional list of model names to restrict seeding to. If None, all seeders defined in the YAML file will be executed.

show_log bool

Flag to enable/disable seeding logs.

seed_stats Any

Stores statistics about the seeding process.

Parameters:

Name Type Description Default
seeders_yml_path str

Path to the seeders YAML configuration file. Defaults to "./app/seed/seeders.yml".

required
data_path str

Path to the default directory containing seed data files.

None
data_paths list[str]

Additional directories containing seed data files. Defaults to "./app/seed/data".

None
models_to_seed list[str]

Names of specific seeders/models to run. If not provided, runs all seeders defined in the configuration.

None
show_log bool

Whether to log seeding operations. Defaults to True.

True

Methods:

Name Description
run

Optional[Engine], session: Optional[Session]): Executes all (or filtered) seeders defined in the YAML configuration using the provided database engine. Supports both: - Custom seeder classes (must subclass SeederBase). - Action-based seeders.

run(engine: Engine | None = None, session: Session | None = None, context: SeedContext | None = None, models_to_seed: list[str] | None = None, reset_context: bool = False, dry_run: bool = False, dry_run_verbose: bool = False, dry_run_quiet: bool = False) -> SeedRunResult

Run all seeders defined in the YAML configuration against the given database engine.

This method
  1. Opens a database session.
  2. Collects seeding statistics.
  3. Iterates over all configured seeders.
  4. For each seeder:
    • Loads associated records from YAML/JSON/CSV files if defined.
    • Validates the configuration schema.
    • Runs either:
      • A custom seeder class (subclass of SeederBase).
      • An action-based seeder.

Parameters:

Name Type Description Default
engine Engine

A SQLAlchemy engine connected to the target database.

None

Raises:

Type Description
AttributeError

If configuration is invalid, or if imports fail (e.g., seeder class/action/model cannot be loaded).

schema

ActionSchema

Bases: SchemaBase

Action seeder specification.

Parameters:

Name Type Description Default
class_name

Optional fully-qualified action class path (e.g., 'app.actions.AccountAction').

required
class_factory

Optional fully-qualified factory function path (e.g., 'app.actions.factories.build_account_action'). If specified, factory is called instead of direct class instantiation.

required

SeederConfig

Bases: SchemaBase

Configuration for running seeders from a YAML file.

Parameters:

Name Type Description Default
seeders_yml_path

Path to the seeders YAML file.

required
data_paths

List of directories to search for seeder data files.

required
models_to_seed

Optional list of specific seeder names to run.

required
show_log

Whether to log per-row seeding operations.

required

utils

format_seeder_validation_error(error, seeder_config, seeders_yml_path: str) -> str

Backward-compatible export; delegates to shared YAML utilities module.

import_lib(module_name: str, module_attr: str)

Dynamically imports a specific attribute (e.g., function, class) from a module.

Parameters:

Name Type Description Default
module_name str

The full dotted path of the module (e.g., 'os.path').

required
module_attr str

The name of the attribute to import from the module.

required

Returns:

Name Type Description
Any

The imported attribute (e.g., class, function, variable) from the specified module.

Raises:

Type Description
ImportError

If the attribute does not exist in the module or the import fails.

matches_filters(model: SQLModel, filters: dict[str, Any]) -> bool

Check if a model instance matches all filter conditions. Args: model (SQLModel): The model instance to check. filters (Dict[str, Any]): A dictionary of field-value pairs to filter by. Returns: bool: True if all conditions match, False otherwise.

get_dependency_value(config: ConfigSchema, seeds: dict[str, Any]) -> dict[str, Any]

Extract dependency values for the current seeder. Args: config (ConfigSchema): The configuration containing dependency details. Returns: Dict[str, Any]: A dictionary containing values to be passed as an update for the current model.

resolve_records_path(records: str, data_paths: list[str] | None = None) -> Path

Resolve a records file path from absolute paths or configured data directories.

resolve_env_placeholders(row: dict[str, Any]) -> dict[str, Any]

Replace values in a dictionary that are in the format ${VAR_NAME} with the value of the corresponding environment variable.

Parameters:

Name Type Description Default
row Dict[str, Any]

A dictionary representing a row of data.

required

Returns:

Type Description
dict[str, Any]

Dict[str, Any]: The updated row with environment variables resolved.

Raises:

Type Description
EnvironmentError

If a referenced environment variable is not set.

normalize_null(value: Any) -> Any | None

Normalize string representations of null values to Python's None.

This function converts specific string values often found in CSVs (e.g., "null", "NULL", "None") into None. All other values are returned unchanged.

Parameters:

Name Type Description Default
value Any

The value to normalize.

required

Returns:

Type Description
Any | None

Optional[Any]: None if the value is a recognized null-like string,

Any | None

otherwise the original value.

get_records(records: str | list[dict[str, Any]] | None = None, data_path: str | None = None, data_paths: list[str] | None = None) -> list[dict[str, Any]]

Load seed data from a CSV/JSON/YAML file or a list of dictionaries, resolving ${VAR_NAME} placeholders from environment variables.

Parameters:

Name Type Description Default
records Optional[Union[str, List[Dict[str, Any]]]]

File name/path or a list of record dictionaries.

None
data_path Optional[str]

Single directory path for record files.

None
data_paths Optional[List[str]]

Ordered list of directories for record files.

None

Returns:

Type Description
list[dict[str, Any]]

List[Dict[str, Any]]: A list of dictionaries for records.

get_yml_config(seeders_yml_path: str) -> dict[str, Any]

Loads and parses a YAML configuration file.

Parameters:

Name Type Description Default
seeders_yml_path str

Path to the YAML file.

required

Returns:

Type Description
dict[str, Any]

Dict[str, Any]: Parsed contents of the YAML file as a dictionary.

Raises:

Type Description
ValueError

If the YAML file contains invalid syntax.

OSError

If the file cannot be read from disk.

TypeError

If the YAML root is not a mapping/object.

validate_path(path: str)

Validates that a given file or directory path exists.

Parameters:

Name Type Description Default
path str

The file or directory path to validate.

required

Returns:

Name Type Description
str

The validated path if it exists.

Raises:

Type Description
ValueError

If the path does not exist.

get_seed_stats(session: Session, logger: logging.Logger, show_log: bool = True)

Attach an event listener to track DB-created instances. Logs every instance that gets inserted during the session.