Skip to content

Seeding

msflib.seed loads initial or demo data from a YAML file. Each entry names a ModelAction, supplies records inline or from a CSV, JSON or YAML file, and can copy values (such as a foreign key) from rows seeded earlier in the same run. It goes through your actions, so validation and hashing behave as in normal code. Lifecycle events go to the current emitter, which is the default emitter unless you wrap the call in with use_app_emitter(app): (see below).

Use it for first-run data, demo data and test fixtures. It is not a migration tool and it is not idempotent (see pitfalls).

Layout

Keep seed code in one folder of your app:

app/
    seed/
        seeders.yml
        data/
            customers.csv
        factories.py
        runner.py

seeders.yml

A top-level seeders mapping. Each entry is a ConfigSchema and needs either action or seeder_class.

seeders:
  customers:
    name: customers
    action:
      class_name: demoapp.actions.CustomerAction
    records: customers.csv

  orders:
    name: orders
    action:
      class_factory: demoapp.seed.factories.order_action_factory
    records:
      - sku: SKU-1
    dependents:
      - name: customers
        foreign_key_field: id
        foreign_key: customer_id
        filter_by:
          name: Grace

Fields:

Field Meaning
name Key under which the created rows are stored for later seeders. Match the key of the mapping.
action.class_name Dotted path to a ModelAction subclass. It is instantiated with no arguments.
action.class_factory Dotted path to a function that returns an action. Use this when the action needs settings or is a shared instance.
seeder_class Dotted path to a SeederBase subclass, for seeding that is not one action.
records A list of dicts, or a file name resolved against the data paths (.csv, .json, .yaml, .yml).
dependents Values to copy from earlier seeders (see below).

Paths must be fully qualified; a name without a dot is rejected.

Each record is passed to action.random(**record), so any required field you leave out is filled with a random value, and the action's create is called with it. Values of the form ${ENV_VAR} in records are replaced from the environment (and raise if unset), and the strings null, NULL and None become None. CSV values arrive as strings, so rely on your create schema to coerce them.

Dependents

A dependents entry copies one attribute of an earlier seeded row into each new row. In the example, orders takes id from the seeded customers row whose name is Grace and stores it as customer_id. Without filter_by the first row of that seeder is used. If the parent seeder produced nothing, or no row matches, the seeder fails.

Factories

A class_factory function receives only the keyword arguments it declares from this set: action_config (also as config), runner, seed_context (also as context), seeds, data_paths, show_log. A function with **kwargs receives all of them.

from demoapp.actions import OrderAction


def order_action_factory(**kwargs):
    return OrderAction()

Custom seeders

Subclass SeederBase when you need more than one create per record. Implement __init__(config) and seed(session, seeds), and store what you create in seeds under config.name so later seeders can depend on it. ActionSeeder in msflib.seed.base is the built-in implementation to copy from.

Running from code

Known issue (#276)

msflib.seed exports only the command-line helpers, so SeedRunner is imported from msflib.seed.runner and SeederBase and ActionSeeder from msflib.seed.base.

SeedRunner runs a file against an engine or an existing session. This example runs against in-memory SQLite, with the Customer and Order models and actions of the YAML above. Import the module that defines the models before create_all; without it no tables are created, both seeders are skipped and the log still ends with "Seeding completed successfully":

from pathlib import Path

from demoapp import models  # noqa: F401
from msflib.seed.runner import SeedRunner
from sqlalchemy.pool import StaticPool
from sqlmodel import SQLModel, create_engine

engine = create_engine(
    "sqlite://", connect_args={"check_same_thread": False}, poolclass=StaticPool
)
SQLModel.metadata.create_all(engine)

seed_dir = Path("demoapp/seed")
runner = SeedRunner(
    str(seed_dir / "seeders.yml"),
    data_paths=[str(seed_dir / "data")],
    show_log=False,
)
result = runner.run(engine=engine)

print(result.seeded_seeders)   # ['customers', 'orders']
print(result.skipped_seeders)  # []
print(result.run_seed_stats)   # {'Customer': 2, 'Order': 1}

show_log=False does not silence the seed summary lines, which are still printed.

Listeners registered on the app emitter, such as the workspaces default-workspace hook, run for seeded rows only if the seeding call is wrapped in with use_app_emitter(app): (from msflib.eventbus).

run also accepts session= instead of engine=, models_to_seed=[...], dry_run=True, and a shared SeedContext so several runs can see each other's rows (reset_context=True clears it). models_to_seed pulls in the seeders named as dependents automatically and runs them in file order. Seeders run in the order they appear in the YAML, so list parents before children.

msflib.db.init_db(engine, metadata, create_tables=False, seeder_config=...) runs the same runner from a SeederConfig. With create_tables=True it first drops every table, so use it for tests and prototypes only.

Command line

run_seed_cli gives your app a seed command with dry-run and safety checks. Put it in app/seed/runner.py:

from pathlib import Path

from app.settings import settings
from msflib.seed import run_seed_cli


def _import_models() -> None:
    from app import models  # noqa: F401


def main() -> None:
    seed_dir = Path(__file__).resolve().parent
    run_seed_cli(
        description="Run app seeders.",
        default_seeders_yml=str(seed_dir / "seeders.yml"),
        default_data_paths=[str(seed_dir / "data")],
        core_settings=settings.scope("CORE"),
        import_models=_import_models,
    )


if __name__ == "__main__":
    main()

Run it with python -m app.seed.runner. The repository's testsite does the same in testsite/app/seed/runner.py (its settings live in app.core.config), and testsite/run_seeder.sh is a thin wrapper that runs it with the repository virtualenv's Python.

Known issue (#290)

The seed CLI (and the initial-data CLI) does not activate the app emitter. Listeners registered on it, such as the workspaces default-workspace hook, do not fire for rows the CLI creates. Wrap your own seeding script in with use_app_emitter(app): if you need them.

Flag Effect
--db-url URL Database to seed. Without it, the URL comes from core_settings (SQLITE_DATABASE_URI when USE_SQLITE, else SQLALCHEMY_DATABASE_URI).
--seeders NAME ... Run only these seeders (plus their dependents).
--seeders-yml PATH, --data-path DIR Override the defaults. --data-path can repeat.
--dry-run Validate the YAML, record files and imports without touching the database. Cannot combine with --init-db.
--init-db Drop and recreate all tables first. Destructive.
--quiet, --verbose Less output; more detailed dry-run output.
--yes Skip the confirmation prompt for sensitive databases.

Safety check: only a SQLite file under ./tmp/ or /tmp/ counts as local. Any other target, including in-memory SQLite and every PostgreSQL URL, is treated as sensitive: the CLI prints a warning, asks you to type SEED, and in a non-interactive shell refuses to run unless you pass --yes.

A dry run is a good CI check that the YAML, data files and import paths still resolve:

python -m app.seed.runner --dry-run --quiet

Pitfalls

  • Not idempotent. The seeder always calls create, so running it twice on the same database inserts the rows twice, or fails on a unique constraint. Seed an empty database, use --init-db in development, or make the data creation idempotent in a custom seeder (check before creating).
  • A failing seeder does not abort the run. The error is logged, the seeder is added to result.skipped_seeders, and later seeders continue. Their dependents then fail too. Check skipped_seeders in scripts and tests rather than assuming success.
  • Seeders run in file order, not dependency order. A child listed before its parent fails with "The parent model (...) cannot be found".
  • All models must be imported before running, so that the tables and relationships exist. That is what import_models is for.
  • Random fill-in. Required fields you omit from records are generated randomly by the action's random(), so fixtures that need stable values should state them.
  • The runner uses after_flush to count rows, so its statistics include every row flushed in the session during seeding, including ones created by event listeners.