compare_modal_execution.py | Runloop runloop.devbox.create(), devbox.cmd.exec(), and file APIs versus Modal modal.App.lookup(), modal.Sandbox.create(), Sandbox.exec(), and Sandbox.filesystem. | Run success, nonzero exit, stdout/stderr, Unicode and binary file, and permission cases. Assert expected results on each service and compare them. |
compare_modal_sessions.py | Runloop devbox.shell(name).exec() versus Modal independent Sandbox.exec() calls or a tested persistent session. | Change directory and environment, then submit dependent and concurrent commands. Verify deliberate ordering and context continuity; test cancellation and one background process. Do not assume independent exec calls share shell state. |
compare_modal_lifecycle.py | Runloop suspend/resume versus Modal lifetime and idle termination followed by replacement. | Write a sentinel, checkpoint, change it again, then trigger planned and unexpected termination. Restore from the externally recorded checkpoint and measure recovery time and accepted data loss. Check cleanup and process restart. |
compare_modal_runtime.py | Runloop blueprint and devbox environment versus Modal Image and selected standard or VM Sandbox runtime. | Check packages, architecture, effective user, working directory, permissions, startup, Docker, systemd, FUSE, and resource behavior needed by the workload. Run only applicable kernel-dependent cases on the chosen runtime. |
compare_modal_network.py | Runloop blueprint network policy and tunnels versus Modal Sandbox networking. | Probe an allowed and blocked outbound destination, authenticated inbound access, and an anonymous inbound request. Require every intended allow or deny decision. |
compare_modal_snapshot.py | Runloop devbox.snapshot_disk() and restore versus Modal Sandbox.snapshot_filesystem(ttl=...) and Sandbox.create(image=...); test directory or memory snapshots only if selected. | Drain writers before the final Runloop snapshot and record its event cursor. Use a large file, symlink, executable mode, database backup, and mounted Volume. Compare root-file manifests and application invariants after source termination; verify Volume recovery separately and retention beyond the agreed recovery window. |
compare_modal_axon.py | Runloop Axon paginated events, SQL APIs, and Broker versus the selected target event journal, database, subscriptions, and Modal Functions or Queues. | Freeze publishers and Broker turns, record the last accepted sequence, export every event page and each SQL table in stable pages, then verify the source event count and sequence range, check rows_read_limit_reached on each SQL page, and compare an independent source row count, hashes, and schema after import. Test reconnect/replay, retries, failed SQL batches, Queue loss, and a worker crash after retrieval. Require no lost accepted events or duplicate external effects. A missing target component leaves this check incomplete. |
compare_modal_benchmark.py | Retained Runloop scenario fixtures, scorer definitions, and results versus Harbor’s Modal environment provider or a custom harness. | Map each source scenario and scorer ID to a versioned task and verifier. Run fixed passing, failing, partial-credit, timeout, and infrastructure-error cases; compare each scorer, weights, aggregate, classification, and artifacts. Missing source definitions leave equivalence unverified. |