Recipes¶
Task-oriented answers. All of them continue the payroll example from Getting started.
Run the same module twice with different parameters¶
Bind the config to the step rather than to the store.
Workflow(
withhold,
Step(report, ReportPolicy(detailed=False), name="summary"),
Step(report, ReportPolicy(detailed=True), name="audit_log"),
)
Without an explicit name, the second occurrence is suffixed automatically (report, report_2). Names
only matter for explain() and on_step, so name them when you'll read them in logs.
Unit-test a module without a workflow¶
A module is callable on a store and returns its operations without applying them. No mocks, no fixtures, no workflow.
def test_overtime_is_paid_at_1_25():
store = Store(
Employee(name="ada", hourly_rate=50.0),
Timesheet(name="ada-w1", employee="ada", hours=38.0),
PayrollPolicy(),
)
(patch,) = compute_gross(store)
assert patch.fields == {"gross": 1937.50}
assert store.find(Employee, "ada").gross == 0.0 # nothing was written
Pass step configs as the second argument when you don't want them in the store:
Run without touching the original store¶
run mutates in place. Copy first when you need the input state afterwards — comparing scenarios,
re-running with different parameters, keeping a before/after.
Compare scenarios¶
No scenario machinery in the library: a loop over configs is a loop over configs.
results = {
rate: Workflow(compute_gross, Step(withhold, PayrollPolicy(social_rate=rate))).run(copy.deepcopy(store))
for rate in (0.20, 0.22, 0.25)
}
{rate: sum(s.net for s in st.all(Payslip)) for rate, st in results.items()}
Log or trace a run¶
To ask the question the other way round — which step wrote this object? — use store.history(obj)
instead of reconstructing it from the log.
ops is the list of Put/Patch/Delete/entity operations the step just wrote — useful for a business log
(f"{len(ops)} {type(ops[0]).__name__}") or for spotting a step that silently did nothing. The same hook
covers progress bars, metrics, and writing intermediate results:
def checkpoint(step: Step, ops: list, store: Store) -> None:
Path(f"out/{step.name}.json").write_text(json.dumps([s.model_dump() for s in store.all(Payslip)]))
workflow.run(store, on_step=checkpoint)
Load a Config from a file¶
Config is a plain pydantic model, so this needs nothing from morphly:
policy = PayrollPolicy.model_validate(tomllib.loads(Path("payroll.toml").read_text()))
store.put(policy)
Same shape for YAML, JSON, environment variables, or a settings service — all validated on the way in.
Give a module its own working fields¶
When a module needs fields nobody else cares about, don't put them on the shared type. Build a local view.
class EmployeeWithHours(Employee):
worked: float
@module
def compute_gross(employees: list[Employee], sheets: list[Timesheet]) -> list[Patch[Employee]]:
hours = {s.employee: s.hours for s in sheets}
rich = [view(EmployeeWithHours, e, worked=hours.get(e.name, 0.0)) for e in employees]
return [Patch(e, gross=e.hourly_rate * e.worked) for e in rich]
EmployeeWithHours is never stored. The Patch still lands on the right Employee, because targets are
resolved against the type in the return annotation.
Write a read-only step¶
Return None. Exports, dashboards, metrics, assertions on the state.
@module
def check_payroll_balances(slips: list[Payslip], employees: list[Employee]) -> None:
total_gross = sum(e.gross for e in employees)
total_slips = sum(s.gross for s in slips)
if abs(total_gross - total_slips) > 0.01:
raise ValueError(f"payroll does not balance: {total_gross} vs {total_slips}")
A read-only module still declares its reads, so check still verifies them.
Fail at startup, before loading any data¶
check only looks at types, so run it against a store built from your schema — no data needed.
def validate_config_at_boot() -> None:
workflow.check(Store(PayrollPolicy(), *(cls(name="_probe") for cls in SEEDED_TYPES)))
Cheaper still: put it in a test. The workflow's shape is static, so a wrong order is a unit-test failure, not a production incident.
Branch, loop, or run workflows in sequence¶
There's no control flow in Workflow because Python already has it.
def run_payroll(store: Store, *, with_bonuses: bool) -> Store:
Workflow(compute_gross).run(store)
if with_bonuses:
Workflow(add_bonus).run(store)
return Workflow(withhold, archive, report).run(store)
Speed up a large run¶
The default deep-copies every injected value. If profiling says that dominates, turn it off:
Modules then receive the stored objects themselves, and mutating an input does change the shared state. Measure first; the isolation guarantee is worth more than most of what it costs.
Skip unchanged steps while iterating in a notebook¶
Editing step 12 and rerunning the whole workflow recomputes steps 1–11 for nothing if their inputs haven't
moved. reuse skips them:
store = loaded_store()
workflow.run(store, record=True)
# ... tweak add_bonus, rerun from scratch ...
store = loaded_store()
workflow.run(store, reuse=workflow.last_run) # compute_gross: skipped, add_bonus onward: runs
record=True is what fills last_run, and it is off by default: recording costs one deep copy of the
store per step, which is worth it here and pure waste in a batch nobody replays.
The cache lives on the Workflow object, in memory, for as long as the process runs. Nothing is persisted
between processes — see Snapshot and restore for that.
Roll back a failed run¶
A step never applies halfway, but the workflow does: if step 4 of 5 raises, steps 1–3 are already written.
atomic=True makes the whole run all-or-nothing, for the price of one deep copy of the store taken before
the first step.
Snapshot and restore¶
For anything else — comparing before and after, keeping several states around — the store is a plain Python object.
For persistence between processes, pickle works. There is no built-in serialization format — see
Non-goals.