Skip to main content
Start with the executable Runloop and Modal comparison script in runloop-examples for command execution and UTF-8 file round trips. The matrix below specifies broader tests to implement for your workload; its filenames are suggested names for those additional checks. The starter does not establish complete migration equivalence. Pin SDK versions and confirm method signatures before extending it.
Use isolated test resources and the same fixed input fixtures on both platforms. Record the Runloop and Modal SDK and CLI versions, image identifiers, runtime type, region, CPU and memory choices, resource IDs, timestamps, and raw output locations. Normalize only provider-specific IDs and timestamps; retain exit codes, stdout, stderr, bytes, permissions, event order, and scorer values. Clean up resources and record any cleanup failure.
  • Lifetime and idle timeout: Test both limits. An active exec, stdin write, or tunnel connection can affect idle behavior; record exactly which activity was present. See the Sandbox lifecycle.
  • Snapshot retention: Test an explicit filesystem snapshot TTL and restoration from its returned Image. Filesystem snapshots exclude mounted Volumes. Confirm expiry behavior before relying on a retention policy. Memory snapshots are Alpha and require separate eligibility and restore testing.
  • Queue loss: Modal Queues do not guarantee persistence. Ensure the journal or database remains the authoritative copy when a queued item disappears or a worker crashes after get().
  • Runtime compatibility: VM Sandboxes are Beta and have different memory and feature constraints from standard Sandboxes. Record which runtime passed each case.

Runner and evidence contract

An implemented script should accept fixture and output-directory arguments and write machine-readable results with versions, resources, cases, assertions, artifacts, and cleanup. Do not write credentials. Exit nonzero when a required assertion or cleanup fails. Preserve raw provider output beside normalized comparisons. Require exact equality for deterministic scores unless a per-scorer tolerance is declared before running. Set a separate sample count and acceptance threshold for stochastic agent tasks. Record the owner, fixture hash, run date, versions, pass or fail, evidence location, and approved exceptions for each case. Cutover requires representative task success, required network checks, state and database integrity, event recovery under queue loss, and scoring agreement. Demonstrate reverse transfer of target writes, pending work, and side effects for lossless rollback. If no reverse path exists, document the accepted post-cutover data-loss window and test recovery to the frozen source state. See the Modal transition guide for the migration sequence.