> ## Documentation Index
> Fetch the complete documentation index at: https://docs.runloop.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Runloop to Tensorlake validation plan

> Side-by-side SDK test specifications for a proposed Tensorlake transition

<Warning>
  Start with the executable [Runloop and Tensorlake comparison script](https://github.com/runloopai/runloop-examples/blob/main/transitions/runloop-to-tensorlake/compare_core_public.py) in `runloop-examples` for command execution and UTF-8 file round trips. The matrix below specifies broader tests to implement for your workload; its filenames are suggested names for those additional checks. The starter does not establish complete migration equivalence. Pin SDK versions and confirm method signatures before extending it.
</Warning>

When implementing these tests, use isolated resources and pinned library versions. Record `runloop` and `tensorlake` package versions, CLI versions, image digest, region, resource sizes, timestamps, resource IDs, and result artifacts. Use the same input fixture on both services. Normalize only provider-specific IDs and timestamps in the raw comparison; record any deliberate application-level output adapter as a separate assertion. Keep exit codes, stdout, stderr, bytes, permissions, and event order intact in raw evidence. Clean up resources after collecting evidence.

| Suggested test filename | Direct library comparison | Fixture and assertions |
| - | - | - |
| `compare_devbox_sandbox.py` | Runloop `runloop.devbox.create()`, `devbox.cmd.exec()` and file APIs versus Tensorlake `Sandbox.create()`, command and file APIs. | Create from pinned environments; run success, nonzero-exit, stdout/stderr, Unicode and binary file cases. Compare raw output, then test an explicit line adapter or byte-preserving file capture if required. Cover empty output, no final newline, multiple final newlines, and failures. Assert exit status, bytes, modes, and ownership. |
| `compare_shell_process.py` | Runloop `devbox.shell(name).exec()` versus Tensorlake persistent PTY or explicit per-command environment and working directory; Runloop asynchronous command versus Tensorlake managed process. | Change directory and export a variable, then submit concurrent dependent commands. Assert ordering and continuity; cancel queued and running work; verify one background process and captured output. |
| `compare_lifecycle.py` | Runloop devbox suspend/resume versus Tensorlake named sandbox suspend/resume; test Tensorlake ephemeral timeout separately. | Start a long-lived counter and write a sentinel. On resume assert disk contents on both, Runloop process restart policy, and no duplicate Tensorlake process. Verify timeout and cleanup states. |
| `compare_network.py` | Runloop blueprint [network policy](/docs/devboxes/blueprints/network-policies) and [tunnels](/docs/devboxes/tunnels) versus Tensorlake [networking](https://docs.tensorlake.ai/sandboxes/networking). | Test an allowed destination, denied destination, authenticated inbound request, and anonymous inbound request. Require each expected allow or deny result. |
| `compare_image.py` | Runloop blueprint provisioning versus Tensorlake registered image provisioning. | Check package versions, user, working directory, startup command, file modes, ports, and absence of embedded runtime secrets. |
| `compare_snapshot.py` | Runloop `devbox.snapshot_disk()` and restore versus Tensorlake `sandbox.checkpoint(checkpoint_type=...)` and `Sandbox.create(snapshot_id=...)`. | Drain writers before the final Runloop snapshot and record its event cursor. Use a fixture with large file, symlink, executable mode, database backup, and external mount. Wait for Tensorlake checkpoint completion before terminating the source; compare manifests and application invariants after restoration. Check filesystem and memory types independently. |
| `compare_axon.py` | Runloop Axon paginated events, SQL APIs, and Broker versus the chosen target journal, subscription, database, and Tensorlake workflow calls. | Freeze publishers and Broker turns, record the last accepted sequence, export every event page and each SQL table in stable pages, then verify the source event count and sequence range, check `rows_read_limit_reached` on each SQL page, and compare an independent source row count, hashes, and schema after import. Test reconnect/replay, retry, concurrent writes, failed SQL batch, and cancellation. Require no lost accepted events or duplicate external effects. A missing target component leaves this check incomplete. |
| `compare_benchmark.py` | Retained Runloop scenario fixtures, scorer definitions, and results versus Harbor or custom harness output on Tensorlake. | Map each source scenario and scorer ID to a versioned task and verifier. Run fixed passing, failing, partial-credit, timeout, and infrastructure-error cases; compare per-scorer values, weights, aggregate, classification, and artifacts. Missing source definitions leave equivalence unverified. |

## Contract for additional tests

Each additional test should accept fixture and output-directory arguments and write a machine-readable JSON result containing `versions`, `resources`, `cases`, `assertions`, `artifacts`, and `cleanup`. Do not write credentials. Exit nonzero on an unmet required assertion or incomplete cleanup. Store raw provider output beside the normalized comparison so differences can be audited.

For deterministic evaluation scores, require exact equality unless a per-scorer numerical tolerance is declared before execution. For stochastic agent tasks, fix the model and sampling parameters, run a declared sample count, and apply a preapproved acceptance threshold. Do not infer equivalence from one run.

## Suggested direct SDK probes

These fragments show candidate library calls only. They omit setup, cleanup, assertions, and error handling; the Runloop fragment also needs an async function. Verify method signatures against the pinned SDKs before turning any row above into an executable script.

```python theme={null}
# Runloop: devbox command and disk snapshot probes
from runloop_api_client import AsyncRunloopSDK

runloop = AsyncRunloopSDK()
devbox = await runloop.devbox.create()
command_result = await devbox.cmd.exec("printf 'migration-probe\\n'")
snapshot = await devbox.snapshot_disk()
```

```python theme={null}
# Tensorlake: sandbox command and typed checkpoint probes
from tensorlake.sandbox import Sandbox, CheckpointType, SnapshotWaitCondition

sandbox = Sandbox.create(name="migration-probe")
command_result = sandbox.run("printf", ["migration-probe\\n"])
snapshot = sandbox.checkpoint(
    checkpoint_type=CheckpointType.FILESYSTEM,
    wait_until=SnapshotWaitCondition.COMPLETED,
)
```

Check the exact Runloop constructor and snapshot result shapes against the [SDK documentation](/docs/tools/sdks) and [snapshot guide](/docs/devboxes/snapshots). Check Tensorlake's [SDK reference](https://docs.tensorlake.ai/sandboxes/sdk-reference), [commands](https://docs.tensorlake.ai/sandboxes/commands), and [checkpoint guide](https://docs.tensorlake.ai/sandboxes/snapshots) before turning the sketches into executable scripts.

## Evidence and cutover gate

For each implemented test, record owner, run date, both versions, fixture hash, raw output location, pass/fail, and exception approval. Cutover requires representative task success, required network checks, state and database integrity, event recovery, and scoring agreement. Demonstrate reverse transfer of target writes, pending work, and side effects for lossless rollback. If no reverse path exists, document the accepted post-cutover data-loss window and test recovery to the frozen source state.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.