Choosing a provider
Choose against the behavior your workload requires, then validate it with the pinned SDKs. Tensorlake named sandboxes can suspend and resume memory and running processes. Modal Sandboxes terminate at their lifetime or idle limit and need explicit state capture and replacement. If you need kernel-dependent tools, test Modal’s Beta VM runtime and the Tensorlake image you intend to use. Existing Modal applications may favor Modal; neither provider supplies an Axon-compatible event journal or a direct migration for Axon SQL. Both can run Harbor tasks. Compare runtime compatibility, networking, cost, and recovery objectives before selecting a target. See the Tensorlake transition guide.Prepare the transition
- Follow Modal setup to create an account, install the SDK, and authenticate. Pin the SDK and CLI versions and confirm feature maturity for the chosen runtime.
- Inventory devboxes, blueprint definitions, snapshots, mounted and external data, Axon SQL databases, events, pending work, Broker integrations, and evaluation history.
- Name an owner and agree recovery-time and recovery-point objectives. Keep Runloop available during validation, with only one authoritative writer at each stage.
- Resolve requirements without a confirmed replacement before planning cutover.
runloop-examples for command execution and UTF-8 file round trips. Extend it with the validation plan for your workload before cutover.
Devboxes to Sandboxes
Choose a Modal App for the integration. Sandboxes created from outside a Modal container require an App; usemodal.Sandbox.create() with an Image and explicit runtime, CPU, memory, lifetime, and optional idle timeout. The standard runtime uses gVisor; VM Sandboxes are Beta. Test Docker, systemd, FUSE, and other kernel-dependent tools on the selected runtime.
Modal documents a five-minute default maximum lifetime, configurable up to 24 hours. An active Sandbox.exec, write to Sandbox stdin, or open TCP tunnel connection counts as activity for idle timeout. Modal uses termination rather than Runloop-style disk-preserving suspend/resume. Checkpoint before planned termination and restore into a replacement; set an accepted data-loss window for unexpected termination. See the Sandbox lifecycle.
Translate resources by runtime. Modal CPU values represent physical cores and memory is in MiB. Standard Sandboxes can burst above scalar minimum requests; VM Sandboxes have fixed memory and elastic CPU and currently lack GPU and reload_volumes support. Check allocation and billing before mapping devbox sizes.
Replace command calls with Sandbox execution and files with the current filesystem API. Runloop named shells preserve directory and environment and serialize calls. Independent Modal exec calls do not establish that behavior; use a tested persistent session or pass context per call and serialize dependent work. Modal’s newer Sandbox backend is becoming the default and does not support the deprecated legacy filesystem methods.
Recreate egress restrictions, inbound access, tunnels, and secrets in provisioning code. Modal networking allows public outbound traffic by default but no inbound connection by default; test both permitted and blocked paths, including unauthenticated inbound requests.
Validate: Run a representative task; compare command exit codes, stdout/stderr, file bytes and permissions, cancellation, background processes, ordering, and working-directory behavior. Exercise allowed and denied networking, kernel-dependent tools, both timeout types, restoration from the externally recorded checkpoint, and cleanup.
Axons to coordination and event infrastructure
Axons provide ordered event history and coordination. Modal Functions can process jobs, but Queues are replicated in memory and persistence is not guaranteed. Retrieval removes an item; a Queue is neither an Axon-compatible journal nor sufficient durable work storage. Inventory event ordering, subscriptions, replay, turn delivery, Broker behavior, and Axon SQL schemas and transaction boundaries. Choose durable storage for accepted events, SQL data, pending work, and completion records. Use Functions for processing where appropriate and Queues only for transient dispatch. Persist work before dispatch, make side effects idempotent, and reconcile unfinished work after crashes or queue loss. Implement subscriptions, replay cursors, and per-Axon turn and SQL serialization separately. For the final export, stop new publishes and SQL writes, drain or stop Broker turns, and record the last accepted event sequence and pending work. The Axon list-events API returns events in descending sequence order. Export every page, preserving each event’s sequence, timestamp, origin, source, type, and payload, then restore ascending order in the target journal. In Runloop Python SDK 1.32.0, passing the last event’s sequence as the stringstarting_after cursor returned the next page in a disposable probe; verify pagination and the source event count with your pinned SDK before cutover.
Export each SQL table in pages of fewer than 1,000 rows, starting each query after the last value of a stable unique indexed key. Fail the export if any query reports rows_read_limit_reached=true. Obtain a source row count separately from the exported pages, then compare it with the imported count and content hashes. Preserve schema and constraints. If a table lacks a usable key or a complete query result, confirm another supported export method before cutover. Keep both stores frozen until both exports and target checks finish. The event stream and SQL database do not share a documented atomic export, so active writers cannot establish one consistent boundary. Test the procedure on a disposable Axon first. See Axon SQL limits.
There is no like-for-like substitute for Axons. Modal is not a durable workflow engine. The recommended approach is to combine Modal Functions, using retries and .spawn(), with your own durable storage to mimic durable workflow execution, and to mirror Axon SQL data to your own SQL instance.
Axon events are immutable, and each Axon holds up to 10GB of event data; size an append-only replacement journal accordingly. A Broker can suspend its devbox when idle and resume it when a new Axon event arrives. Modal has no disk-preserving suspend, so replace this with a process that starts a new Sandbox, restored from a snapshot, when work arrives.
Validate: Publish ordered events; disconnect and replay; cancel and interrupt turns; submit concurrent messages; lose or expire queued items; and crash a worker immediately after retrieval. Require no lost accepted events or duplicate external effects. Verify migrated SQL schemas, values, constraints, serialized operations, and all-or-nothing rollback of failed batches.
Blueprints to Images
Record each blueprint base, dependencies, files, setup, user, and working directory. Translate it into a Modal Image, supported Dockerfile, or registry image. External images must targetlinux/amd64. Modal ignores Dockerfile USER and runs containers as root; explicitly drop process privileges if needed. Some Dockerfile instructions, including HEALTHCHECK and VOLUME, are not supported. Review image compatibility.
Keep runtime credentials separate from reusable images. Reapply resource choices, network policy, ports, mounts, and secrets when creating Sandboxes. Record the blueprint-to-image mapping.
Validate: In a fresh Sandbox, compare packages, architecture, paths, effective user, permissions, startup behavior, agent version, and access controls. Confirm runtime secrets are absent from the image.
Snapshots to Sandbox snapshots
Direct Runloop snapshot import into Modal is unverified. For the final cutover, stop new work, drain or record in-flight work, quiesce filesystem and database writers, flush writes, and record the event cursor before creating the Runloop snapshot. Wait for completion, restore it into a temporary devbox, and export required application directories plus an application-consistent database backup. Verify that no accepted writes occurred after the recorded boundary; if work could not be drained, capture and reconcile the final delta before enabling the target writer. Preserve checksums, modes, and symlinks; handle mounts and external data separately. Transfer through supported authenticated file workflows, map ownership, restore into the new Sandbox, and verify application invariants before making a Modal snapshot. A Modal filesystem snapshot produces an Image of the root filesystem for new Sandboxes; mounted Volumes are excluded. Directory snapshots capture a selected directory. For Python SDK 1.5+ and Go/JS SDK 0.8.0+, filesystem snapshots default to 30-day retention; directory snapshots also default to 30 days. In Python, passttl=<seconds> or ttl=None deliberately. Store the Image reference, event cursor, and pending-work reconciliation metadata outside the Sandbox.
Memory snapshots are Alpha, expire after seven days, and cannot be extended. They terminate the source Sandbox, cannot be taken during an active Sandbox.exec, do not properly restore background processes started through exec, close TCP connections, and cannot use GPUs. Restores require the same instance type. VM memory snapshots have separate customer gating. Use them only after confirming eligibility and testing their constraints.
Validate: Compare file manifests, large files, symlinks, permissions, database integrity, and event cursors. Restore after source termination, restart required processes, reattach and check Volumes, inject capture and recovery failures, and verify retention covers the recovery objective. Check that reusable snapshots do not include unintended credentials or customer data.
Benchmarks to Harbor
The Runloop Benchmark product has been discontinued. The recommended migration path is to switch to the Harbor runner and select Modal as the sandbox provider withharbor run --env modal; see Harbor on Modal. Harbor’s pre-integrated sandboxes page has more details and lists alternative providers.
Treat this as a task and scoring port. Inventory each retained scenario’s instruction, input fixture, environment, agent version, timeout, scorer code and parameters, weights, partial-credit rules, failure classification, and historical artifacts. Map each scenario to a Harbor task with instruction.md, task.toml, environment configuration, and verifier code under tests/. Port the original scoring contract to Harbor’s verifier reward output; RewardKit can represent multiple weighted criteria. Keep a versioned mapping from Runloop scenario and scorer IDs to Harbor tasks and verifiers. Compare fixed passing, failing, and partial-credit fixtures per scorer before aggregate scores. Archive historical Runloop results separately unless a supported import path is confirmed. If original fixtures or scorer definitions are unavailable, mark score equivalence unverified.
