MLExperiment Manager · for macOS

Your training repo, improving itself.

Import a repo and an agent reads the code, maps every lever, and proposes experiments. Each one runs across seeds, is judged against the baseline's noise, and the winners merge back.

Propose. Run. Judge. Merge.

  1. 1

    Import a repo

    An agent reads the code, maps its levers, and writes a compatibility patch, so the repo reports metrics, saves checkpoints and can be scored without touching your branch.

  2. 2

    Run across seeds

    On this Mac or in a Modal GPU sandbox. Failed runs are diagnosed and repaired by the agent, then relaunched.

  3. 3

    Judged against the noise

    A proposal wins only when its gain is larger than the baseline's seed noise. One lucky seed is not a result.

  4. 4

    Branch from what works

    Grow a tree of experiments, talk new ones through with Claude, combine two branches, and merge the winners into your default branch.

Everything else the loop needs.

Evals per checkpoint

Every registered eval set scores every checkpoint, outside the run's time budget.

Sample galleries

Each checkpoint is sampled with one fixed request, so runs line up side by side.

Playground

Sample any kept checkpoint with your own prompt, input image and seed.

Compare

Curves of every seed on one axis, scores per eval set, and an agent's read of the samples.

Modal GPUs

Runs in a GPU sandbox that always terminates. The best checkpoints stay on a Volume.

Cost cap

A run estimated above the cap is refused; one that crosses it is stopped.

Get the app

For macOS on Apple silicon and Intel. Works with Claude Code or OpenRouter, on this Mac or on Modal.

Not notarized yet, so macOS blocks the first launch: open System Settings → Privacy & Security and choose Open Anyway, or run xattr -dr com.apple.quarantine "/Applications/MLExperiment Manager.app".