msflib.seed¶
Part of the msflib core package.
msflib.seed
¶
is_production_db(db_url: str) -> bool
¶
Whether a database URL must be treated as sensitive (conservative heuristic).
Only a SQLite file under ./tmp/ or /tmp/ is local. The path is normalised before it is
compared, so sqlite:///./tmp/../prod.db (which SQLite resolves to ./prod.db) is not
mistaken for a temp file, a path containing a backslash is never local, and symlinks are
followed: a link inside the temp directory that points elsewhere makes the file sensitive.
Everything else, including in-memory SQLite and every non-SQLite URL (also PostgreSQL on
localhost), is sensitive.
run_seed_cli(*, argv: Sequence[str] | None = None, description: str = 'Run database seeders.', default_seeders_yml: str | None = None, default_data_paths: list[str] | None = None, core_settings: CoreSettings | None = None, import_models: Callable[[], None] | None = None, metadata=SQLModel.metadata, is_production_db: Callable[[str], bool] = is_production_db) -> None
¶
Run database seeders from a YAML configuration file.
This is the main entry point for seeding. It: 1. Parses command-line arguments 2. Confirms sensitive database operations 3. Optionally initializes the database schema 4. Runs seeders from the YAML configuration
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
argv
|
Sequence[str] | None
|
Command-line arguments (defaults to sys.argv). |
None
|
description
|
str
|
Description for the CLI help text. |
'Run database seeders.'
|
default_seeders_yml
|
str | None
|
Default path to seeders.yml file. |
None
|
default_data_paths
|
list[str] | None
|
Default list of data directory paths. |
None
|
core_settings
|
CoreSettings | None
|
Optional settings object that provides database URL if not provided via CLI. |
None
|
import_models
|
Callable[[], None] | None
|
Optional callable to import all SQLModel models before initialization. |
None
|
metadata
|
SQLAlchemy MetaData instance (typically SQLModel.metadata). |
metadata
|
|
is_production_db
|
Callable[[str], bool]
|
Function to determine if a database URL is likely in production. |
is_production_db
|
base
¶
SeederBase(config: ConfigSchema)
¶
Bases: ABC
Abstract base class for creating seeders.
This class defines the required structure for any seeder implementation.
Subclasses must implement both the __init__ and seed methods to handle
configuration and seeding logic respectively.
Initialize the seeder with a configuration schema.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
ConfigSchema
|
Configuration object containing parameters needed for seeding. |
required |
Example: >>> # You should initialize your config and records as below: >>> self.config = config >>> self.records = ( ... config.records if config.records else [model_action.random().model_dump()] ... )
seed(session: Session, seeds: dict[str, Any])
abstractmethod
¶
Perform seeding logic using the provided database session and seed data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
session
|
Session
|
SQLAlchemy session object to interact with the database. |
required |
seeds
|
Dict[str, Any]
|
Dictionary of seed data to insert into the database. |
required |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If the method is not implemented by a subclass. |
Example:
>>> # Always update seeds with the seeded instance. for example:
>>> accounts = seeds["accounts"] # getting account instances from the seeds dict
>>> profiles = []
>>> for account, record in zip(accounts, self.records):
profile = pa.create(
session, data=pa.random(**record), update={"id": account.id},
)
profiles.append(profile)
>>> # Storing the profile instances so it can be used as a dependency just like
>>> # how account instances was used to seed profile.
>>> seeds.update({self.config.name: profiles})
ActionSeeder(model_action: ModelActionType, config: ConfigSchema)
¶
Bases: SeederBase
Base class for seeding database tables. Handles: - Fetching dependencies from previously seeded data - Reading data from CSV files - Creating records in the database
Initialize the seeder with the required configurations. Args: model_action (ModelActionType): The action class responsible for creating models. config (ConfigSchema): Configuration containing seeder details.
seed(session: Session, seeds: dict[str, Any])
¶
Seed the database with records. Args: session (Session): The database session for committing transactions. seeds (Dict): This stores the created seeds incase of dependency between seeds.
cli
¶
is_production_db(db_url: str) -> bool
¶
Whether a database URL must be treated as sensitive (conservative heuristic).
Only a SQLite file under ./tmp/ or /tmp/ is local. The path is normalised before it is
compared, so sqlite:///./tmp/../prod.db (which SQLite resolves to ./prod.db) is not
mistaken for a temp file, a path containing a backslash is never local, and symlinks are
followed: a link inside the temp directory that points elsewhere makes the file sensitive.
Everything else, including in-memory SQLite and every non-SQLite URL (also PostgreSQL on
localhost), is sensitive.
run_seed_cli(*, argv: Sequence[str] | None = None, description: str = 'Run database seeders.', default_seeders_yml: str | None = None, default_data_paths: list[str] | None = None, core_settings: CoreSettings | None = None, import_models: Callable[[], None] | None = None, metadata=SQLModel.metadata, is_production_db: Callable[[str], bool] = is_production_db) -> None
¶
Run database seeders from a YAML configuration file.
This is the main entry point for seeding. It: 1. Parses command-line arguments 2. Confirms sensitive database operations 3. Optionally initializes the database schema 4. Runs seeders from the YAML configuration
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
argv
|
Sequence[str] | None
|
Command-line arguments (defaults to sys.argv). |
None
|
description
|
str
|
Description for the CLI help text. |
'Run database seeders.'
|
default_seeders_yml
|
str | None
|
Default path to seeders.yml file. |
None
|
default_data_paths
|
list[str] | None
|
Default list of data directory paths. |
None
|
core_settings
|
CoreSettings | None
|
Optional settings object that provides database URL if not provided via CLI. |
None
|
import_models
|
Callable[[], None] | None
|
Optional callable to import all SQLModel models before initialization. |
None
|
metadata
|
SQLAlchemy MetaData instance (typically SQLModel.metadata). |
metadata
|
|
is_production_db
|
Callable[[str], bool]
|
Function to determine if a database URL is likely in production. |
is_production_db
|
runner
¶
SeedContext(seeds: dict[str, list[ModelBase]] = dict(), skipped_seeders: list[str] = list(), seed_stats: Counter = Counter())
dataclass
¶
Stores seeded entities and cumulative run metadata for repeated seed runs.
SeedRunner(seeders_yml_path: str, data_path: str | None = None, data_paths: list[str] | None = None, models_to_seed: list[str] | None = None, show_log: bool = True)
¶
Handles database seeding by running seeders defined in a YAML configuration file.
The SeedRunner loads seeder configurations, processes record data, dynamically imports custom seeder classes, actions, and models, and executes seeding operations into the database. It also allows selective execution of seeders for specific models.
Attributes:
| Name | Type | Description |
|---|---|---|
context |
SeedContext
|
Shared context storing seeds and cumulative stats. |
seeders |
dict
|
Seeder configurations parsed from the seeders YAML file. |
data_paths |
list[str]
|
Directory paths where seeder data files are searched. |
models_to_seed |
list[str]
|
Optional list of model names to restrict seeding to. If None, all seeders defined in the YAML file will be executed. |
show_log |
bool
|
Flag to enable/disable seeding logs. |
seed_stats |
Any
|
Stores statistics about the seeding process. |
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seeders_yml_path
|
str
|
Path to the seeders YAML configuration file. Defaults to "./app/seed/seeders.yml". |
required |
data_path
|
str
|
Path to the default directory containing seed data files. |
None
|
data_paths
|
list[str]
|
Additional directories containing seed data files. Defaults to "./app/seed/data". |
None
|
models_to_seed
|
list[str]
|
Names of specific seeders/models to run. If not provided, runs all seeders defined in the configuration. |
None
|
show_log
|
bool
|
Whether to log seeding operations. Defaults to True. |
True
|
Methods:
| Name | Description |
|---|---|
run |
Optional[Engine], session: Optional[Session]): Executes all (or filtered) seeders defined in the YAML configuration using the provided database engine. Supports both: - Custom seeder classes (must subclass SeederBase). - Action-based seeders. |
run(engine: Engine | None = None, session: Session | None = None, context: SeedContext | None = None, models_to_seed: list[str] | None = None, reset_context: bool = False, dry_run: bool = False, dry_run_verbose: bool = False, dry_run_quiet: bool = False) -> SeedRunResult
¶
Run all seeders defined in the YAML configuration against the given database engine.
This method
- Opens a database session.
- Collects seeding statistics.
- Iterates over all configured seeders.
- For each seeder:
- Loads associated records from YAML/JSON/CSV files if defined.
- Validates the configuration schema.
- Runs either:
- A custom seeder class (subclass of SeederBase).
- An action-based seeder.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
engine
|
Engine
|
A SQLAlchemy engine connected to the target database. |
None
|
Raises:
| Type | Description |
|---|---|
AttributeError
|
If configuration is invalid, or if imports fail (e.g., seeder class/action/model cannot be loaded). |
schema
¶
ActionSchema
¶
Bases: SchemaBase
Action seeder specification.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
class_name
|
Optional fully-qualified action class path (e.g., 'app.actions.AccountAction'). |
required | |
class_factory
|
Optional fully-qualified factory function path (e.g., 'app.actions.factories.build_account_action'). If specified, factory is called instead of direct class instantiation. |
required |
SeederConfig
¶
Bases: SchemaBase
Configuration for running seeders from a YAML file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seeders_yml_path
|
Path to the seeders YAML file. |
required | |
data_paths
|
List of directories to search for seeder data files. |
required | |
models_to_seed
|
Optional list of specific seeder names to run. |
required | |
show_log
|
Whether to log per-row seeding operations. |
required |
utils
¶
format_seeder_validation_error(error, seeder_config, seeders_yml_path: str) -> str
¶
Backward-compatible export; delegates to shared YAML utilities module.
import_lib(module_name: str, module_attr: str)
¶
Dynamically imports a specific attribute (e.g., function, class) from a module.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
module_name
|
str
|
The full dotted path of the module (e.g., 'os.path'). |
required |
module_attr
|
str
|
The name of the attribute to import from the module. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Any |
The imported attribute (e.g., class, function, variable) from the specified module. |
Raises:
| Type | Description |
|---|---|
ImportError
|
If the attribute does not exist in the module or the import fails. |
matches_filters(model: SQLModel, filters: dict[str, Any]) -> bool
¶
Check if a model instance matches all filter conditions. Args: model (SQLModel): The model instance to check. filters (Dict[str, Any]): A dictionary of field-value pairs to filter by. Returns: bool: True if all conditions match, False otherwise.
get_dependency_value(config: ConfigSchema, seeds: dict[str, Any]) -> dict[str, Any]
¶
Extract dependency values for the current seeder. Args: config (ConfigSchema): The configuration containing dependency details. Returns: Dict[str, Any]: A dictionary containing values to be passed as an update for the current model.
resolve_records_path(records: str, data_paths: list[str] | None = None) -> Path
¶
Resolve a records file path from absolute paths or configured data directories.
resolve_env_placeholders(row: dict[str, Any]) -> dict[str, Any]
¶
Replace values in a dictionary that are in the format ${VAR_NAME} with the value of the corresponding environment variable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
row
|
Dict[str, Any]
|
A dictionary representing a row of data. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict[str, Any]: The updated row with environment variables resolved. |
Raises:
| Type | Description |
|---|---|
EnvironmentError
|
If a referenced environment variable is not set. |
normalize_null(value: Any) -> Any | None
¶
Normalize string representations of null values to Python's None.
This function converts specific string values often found in CSVs
(e.g., "null", "NULL", "None") into None. All other values are
returned unchanged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
The value to normalize. |
required |
Returns:
| Type | Description |
|---|---|
Any | None
|
Optional[Any]: |
Any | None
|
otherwise the original value. |
get_records(records: str | list[dict[str, Any]] | None = None, data_path: str | None = None, data_paths: list[str] | None = None) -> list[dict[str, Any]]
¶
Load seed data from a CSV/JSON/YAML file or a list of dictionaries, resolving ${VAR_NAME} placeholders from environment variables.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
records
|
Optional[Union[str, List[Dict[str, Any]]]]
|
File name/path or a list of record dictionaries. |
None
|
data_path
|
Optional[str]
|
Single directory path for record files. |
None
|
data_paths
|
Optional[List[str]]
|
Ordered list of directories for record files. |
None
|
Returns:
| Type | Description |
|---|---|
list[dict[str, Any]]
|
List[Dict[str, Any]]: A list of dictionaries for records. |
get_yml_config(seeders_yml_path: str) -> dict[str, Any]
¶
Loads and parses a YAML configuration file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seeders_yml_path
|
str
|
Path to the YAML file. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict[str, Any]: Parsed contents of the YAML file as a dictionary. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the YAML file contains invalid syntax. |
OSError
|
If the file cannot be read from disk. |
TypeError
|
If the YAML root is not a mapping/object. |
validate_path(path: str)
¶
Validates that a given file or directory path exists.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str
|
The file or directory path to validate. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The validated path if it exists. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the path does not exist. |
get_seed_stats(session: Session, logger: logging.Logger, show_log: bool = True)
¶
Attach an event listener to track DB-created instances. Logs every instance that gets inserted during the session.