> ## Documentation Index
> Fetch the complete documentation index at: https://docs.runloop.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Runloop to Modal validation plan

> Side-by-side SDK test specifications for a proposed Modal transition

<Warning>
  Start with the executable [Runloop and Modal comparison script](https://github.com/runloopai/runloop-examples/blob/main/transitions/runloop-to-modal/compare_core_public.py) in `runloop-examples` for command execution and UTF-8 file round trips. The matrix below specifies broader tests to implement for your workload; its filenames are suggested names for those additional checks. The starter does not establish complete migration equivalence. Pin SDK versions and confirm method signatures before extending it.
</Warning>

Use isolated test resources and the same fixed input fixtures on both platforms. Record the Runloop and Modal SDK and CLI versions, image identifiers, runtime type, region, CPU and memory choices, resource IDs, timestamps, and raw output locations. Normalize only provider-specific IDs and timestamps; retain exit codes, stdout, stderr, bytes, permissions, event order, and scorer values. Clean up resources and record any cleanup failure.

| Suggested test filename | Direct library comparison | Fixture and assertions |
| - | - | - |
| `compare_modal_execution.py` | Runloop `runloop.devbox.create()`, `devbox.cmd.exec()`, and file APIs versus Modal `modal.App.lookup()`, `modal.Sandbox.create()`, `Sandbox.exec()`, and `Sandbox.filesystem`. | Run success, nonzero exit, stdout/stderr, Unicode and binary file, and permission cases. Assert expected results on each service and compare them. |
| `compare_modal_sessions.py` | Runloop `devbox.shell(name).exec()` versus Modal independent `Sandbox.exec()` calls or a tested persistent session. | Change directory and environment, then submit dependent and concurrent commands. Verify deliberate ordering and context continuity; test cancellation and one background process. Do not assume independent exec calls share shell state. |
| `compare_modal_lifecycle.py` | Runloop suspend/resume versus Modal lifetime and idle termination followed by replacement. | Write a sentinel, checkpoint, change it again, then trigger planned and unexpected termination. Restore from the externally recorded checkpoint and measure recovery time and accepted data loss. Check cleanup and process restart. |
| `compare_modal_runtime.py` | Runloop blueprint and devbox environment versus Modal Image and selected standard or VM Sandbox runtime. | Check packages, architecture, effective user, working directory, permissions, startup, Docker, systemd, FUSE, and resource behavior needed by the workload. Run only applicable kernel-dependent cases on the chosen runtime. |
| `compare_modal_network.py` | Runloop blueprint [network policy](/docs/devboxes/blueprints/network-policies) and [tunnels](/docs/devboxes/tunnels) versus Modal [Sandbox networking](https://modal.com/docs/guide/sandbox-networking). | Probe an allowed and blocked outbound destination, authenticated inbound access, and an anonymous inbound request. Require every intended allow or deny decision. |
| `compare_modal_snapshot.py` | Runloop `devbox.snapshot_disk()` and restore versus Modal `Sandbox.snapshot_filesystem(ttl=...)` and `Sandbox.create(image=...)`; test directory or memory snapshots only if selected. | Drain writers before the final Runloop snapshot and record its event cursor. Use a large file, symlink, executable mode, database backup, and mounted Volume. Compare root-file manifests and application invariants after source termination; verify Volume recovery separately and retention beyond the agreed recovery window. |
| `compare_modal_axon.py` | Runloop Axon paginated events, SQL APIs, and Broker versus the selected target event journal, database, subscriptions, and Modal Functions or Queues. | Freeze publishers and Broker turns, record the last accepted sequence, export every event page and each SQL table in stable pages, then verify the source event count and sequence range, check `rows_read_limit_reached` on each SQL page, and compare an independent source row count, hashes, and schema after import. Test reconnect/replay, retries, failed SQL batches, Queue loss, and a worker crash after retrieval. Require no lost accepted events or duplicate external effects. A missing target component leaves this check incomplete. |
| `compare_modal_benchmark.py` | Retained Runloop scenario fixtures, scorer definitions, and results versus Harbor's Modal environment provider or a custom harness. | Map each source scenario and scorer ID to a versioned task and verifier. Run fixed passing, failing, partial-credit, timeout, and infrastructure-error cases; compare each scorer, weights, aggregate, classification, and artifacts. Missing source definitions leave equivalence unverified. |

## Modal-specific gates

* **Lifetime and idle timeout:** Test both limits. An active exec, stdin write, or tunnel connection can affect idle behavior; record exactly which activity was present. See the [Sandbox lifecycle](https://modal.com/docs/guide/sandboxes).
* **Snapshot retention:** Test an explicit filesystem snapshot TTL and restoration from its returned Image. Filesystem snapshots exclude mounted Volumes. Confirm expiry behavior before relying on a retention policy. [Memory snapshots](https://modal.com/docs/guide/sandbox-snapshots) are Alpha and require separate eligibility and restore testing.
* **Queue loss:** [Modal Queues](https://modal.com/docs/guide/queues) do not guarantee persistence. Ensure the journal or database remains the authoritative copy when a queued item disappears or a worker crashes after `get()`.
* **Runtime compatibility:** [VM Sandboxes](https://modal.com/docs/guide/vm-sandboxes) are Beta and have different memory and feature constraints from standard Sandboxes. Record which runtime passed each case.

## Runner and evidence contract

An implemented script should accept fixture and output-directory arguments and write machine-readable results with `versions`, `resources`, `cases`, `assertions`, `artifacts`, and `cleanup`. Do not write credentials. Exit nonzero when a required assertion or cleanup fails. Preserve raw provider output beside normalized comparisons.

Require exact equality for deterministic scores unless a per-scorer tolerance is declared before running. Set a separate sample count and acceptance threshold for stochastic agent tasks. Record the owner, fixture hash, run date, versions, pass or fail, evidence location, and approved exceptions for each case.

Cutover requires representative task success, required network checks, state and database integrity, event recovery under queue loss, and scoring agreement. Demonstrate reverse transfer of target writes, pending work, and side effects for lossless rollback. If no reverse path exists, document the accepted post-cutover data-loss window and test recovery to the frozen source state. See the [Modal transition guide](/docs/transitions/runloop-to-modal) for the migration sequence.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.