Skip to main content
This is a planning guide based on published documentation, not a certified migration. Pin SDK and CLI versions, run the side-by-side validation plan, and record evidence for your workload before cutover.
Runloop devboxes, blueprints, and snapshots have related Tensorlake resources, but resource IDs and stored state do not transfer automatically. Axons require application-level design work, and the discontinued Benchmark product is replaced by the Harbor runner.

Choosing a provider

Choose against the behavior your workload requires, then validate it with the pinned SDKs. Tensorlake named sandboxes can suspend and resume memory and running processes. Modal Sandboxes terminate at their lifetime or idle limit and need explicit state capture and replacement. If you need kernel-dependent tools, test Modal’s Beta VM runtime and the Tensorlake image you intend to use. Existing Modal applications may favor Modal; neither provider supplies an Axon-compatible event journal or a direct migration for Axon SQL. Both can run Harbor tasks. Compare runtime compatibility, networking, cost, and recovery objectives before selecting a target. See the Modal transition guide.

Prepare the transition

  1. Inventory devboxes, blueprint definitions, snapshot IDs, Axon event and SQL data, Broker usage, evaluation fixtures, mounted and external data, and historical results.
  2. Choose one authoritative writer for each stage. Keep the Runloop integration available for recovery, without allowing both systems to accept independent writes.
  3. Pin both vendors’ SDK and CLI versions, name an owner, and agree recovery-time and recovery-point objectives. Record the test output and exceptions.
  4. Set a project-scoped TENSORLAKE_API_KEY for SDK processes. Tensorlake CLI login is a separate authentication path; follow its quickstart.
Start with the Runloop and Tensorlake SDK comparison script in runloop-examples for command execution and UTF-8 file round trips. Extend it with the validation plan for your workload before cutover.

Devboxes to sandboxes

Use sandbox creation, commands and processes, and file operations in place of devbox calls. Set CPU, memory, disk, image, and timeout for each workload. Sandbox resources are fixed after creation, so resizing requires a new sandbox. Only named sandboxes suspend on idle; ephemeral sandboxes terminate. timeout_secs measures idle time without in-flight proxy traffic. The documented default is 600 seconds, and 0 requests the plan maximum. SSH, PTY, exposed-port traffic, and SDK or CLI calls can reset idleness. A terminated sandbox may be restarted within 48 hours, but termination is not a guarantee of current filesystem state. See the Tensorlake lifecycle guide. Runloop suspend/resume preserves disk state but requires processes to be restarted. Tensorlake named sandbox suspension preserves memory and running processes. Remove unconditional process startup on resume or protect it against duplicates. Runloop named shells serialize calls and preserve directory and environment; verify equivalent behavior with a persistent Tensorlake session or pass context on every command and serialize calls in your application. Command output compatibility: In a probe with Runloop Python SDK 1.32.0 and Tensorlake Python SDK 0.5.135, Sandbox.run() returned the same output lines but omitted trailing newlines present in Runloop’s stdout and stderr. If callers consume lines, define and test an explicit line-based adapter. If they require exact bytes, redirect stdout and stderr to separate sandbox files and read those files as bytes while retaining the command’s exit status. Test empty output, missing or multiple final newlines, binary output, and failures. Do not add or strip a newline unconditionally. Tensorlake documents command redirection and file reads. Recreate outbound policy and inbound authentication deliberately. Tensorlake documents outbound access as enabled by default and distinguishes authenticated access from public port exposure in its networking guide. Validate allowed and denied routes, including anonymous inbound requests.

Axons to coordination and event infrastructure

Axons provide ordered events, observation, and coordination. Tensorlake Orchestration offers durable workflow execution and retries, but its documentation does not establish an Axon-compatible event journal, replay stream, subscription API, or Broker protocol. Define each required behavior before choosing replacement components. There is no like-for-like substitute for Axons. The recommended approach is to use Tensorlake Orchestration to mimic durable workflow execution and to mirror Axon SQL data to your own SQL instance. Migrate Axon SQL as its own dataset, separate from event history and devbox disk. Inventory schemas, transactions, constraints, serialized access, and atomic batches. Specify a durable event journal and subscriber recovery cursor if clients need replay. Make workflow side effects safe to retry, and replace Broker turn delivery explicitly. For the final export, stop new publishes and SQL writes, drain or stop Broker turns, and record the last accepted event sequence and pending work. The Axon list-events API returns events in descending sequence order. Export every page, preserving each event’s sequence, timestamp, origin, source, type, and payload, then restore ascending order in the target journal. In Runloop Python SDK 1.32.0, passing the last event’s sequence as the string starting_after cursor returned the next page in a disposable probe; verify pagination and the source event count with your pinned SDK before cutover. Export each SQL table in pages of fewer than 1,000 rows, starting each query after the last value of a stable unique indexed key. Fail the export if any query reports rows_read_limit_reached=true. Obtain a source row count separately from the exported pages, then compare it with the imported count and content hashes. Preserve schema and constraints. If a table lacks a usable key or a complete query result, confirm another supported export method before cutover. Keep both stores frozen until both exports and target checks finish. The event stream and SQL database do not share a documented atomic export, so active writers cannot establish one consistent boundary. Test the procedure on a disposable Axon first. See Axon SQL limits.

Blueprints to sandbox images

Record each blueprint’s base image, packages, files, user, working directory, and setup commands. Use Tensorlake’s build or import workflow and pass its registered image name to sandbox creation. An unregistered registry reference is not a launchable image. Built and imported images inherit their image USER and WORKDIR, which may differ from Tensorlake managed-image defaults. Keep runtime credentials out of reusable images. Reapply egress policy, inbound authentication, ports, tunnels, and mounts during provisioning; the image alone does not transfer those settings. Map each blueprint ID to its replacement registered image name and test a fresh environment.

Snapshots to checkpoints

Direct Runloop snapshot import into Tensorlake is unverified. For the final cutover, stop new work, drain or record in-flight work, quiesce filesystem and database writers, flush writes, and record the event cursor before creating the Runloop snapshot. Wait for completion, restore it to a temporary devbox, and export only required application directories and a consistent database backup. Preserve a checksum, mode, and symlink manifest. Verify that no accepted writes occurred after the recorded boundary; if work could not be drained, capture and reconcile the final delta before enabling the target writer. Transfer through supported authenticated file workflows, restore into the rebuilt sandbox, map ownership, and verify application invariants. Treat mounted and external data separately. Tensorlake checkpoints distinguish filesystem checkpoints, which cold boot on restore, from memory checkpoints, which also preserve processes. Set the checkpoint type explicitly and verify the behavior for your pinned SDK. When restoring in a new sandbox after terminating the source, wait for SnapshotWaitCondition.COMPLETED before termination; the default local-ready wait was insufficient in our filesystem restore probe with Tensorlake Python SDK 0.5.135. Store the resulting snapshot ID, event cursor, and pending-work reconciliation metadata outside the sandbox. A sandbox checkpoint does not make external side effects or shared-volume writes atomic.

Benchmarks to Harbor

The Runloop Benchmark product has been discontinued. The recommended migration path is to switch to the Harbor runner and select Tensorlake as the sandbox provider; see Harbor on Tensorlake sandboxes. Harbor’s pre-integrated sandboxes page has more details and lists alternative providers. Treat this as a task and scoring port. Inventory each retained scenario’s instruction, input fixture, environment, agent version, timeout, scorer code and parameters, weights, partial-credit rules, failure classification, and historical artifacts. Map each scenario to a Harbor task with instruction.md, task.toml, environment configuration, and verifier code under tests/. Port the original scoring contract to Harbor’s verifier reward output; RewardKit can represent multiple weighted criteria. Keep a versioned mapping from Runloop scenario and scorer IDs to Harbor tasks and verifiers. Compare fixed passing, failing, and partial-credit fixtures per scorer before aggregate scores. Archive historical Runloop results separately unless a supported import path is confirmed. If original fixtures or scorer definitions are unavailable, mark score equivalence unverified.

Cutover and recovery

Run the validation matrix for every primitive in use. Require no unexplained file-manifest differences, no lost accepted events or duplicate external effects under injected failures, all required permission checks passing, and recovery within the agreed objectives. Document any accepted exceptions. Rehearse a final drain, consistent export, target restore, routing change, and acceptance check. Lossless rollback after target writes requires a tested reverse path for new events, SQL changes, files, pending work, and external side effects. Stop target writers, export changes since the cutover boundary, apply them once to Runloop, reconcile side effects, and only then restore routing. Republished Axon events receive new sequence numbers; map target IDs and cursors to source IDs using stable application-level identifiers. Without this tested reverse path, recovery returns only to the frozen pre-cutover state and may lose later writes. State that accepted data-loss window before cutover; consider repairing the target instead of switching back. Runloop is winding down its services. Transition support for existing customers continues through November 13, 2026 at 11:59 PM PT; contact Runloop support for help. Open technical decisions include Axon replacement, snapshot portability, and automated migration tooling.