Closed beta: top-ups come with 30% extra credits, subscriptions with 50–70% extra. We welcome your feedback on the support page.

Measured results

56 open-source projects,
every number verifiable

We ran AutoOptm on these public PyTorch and Python repositories exactly as we handle any submission: the project's own command (or, for a library with no entry point, a benchmark script exercising its public API), timed end to end, with outputs compared item by item before and after.

56public repositories
2.48×median speedup
1.18–23.41×range
59 / 59measurements, output verified

All optimized forks on GitHub↗

14.65×
CPU data processingnoise floor 0.6%

Unit timed: one random graph through six analyses (betweenness, PageRank, clustering, …)
A library with no entry point: timing covers a benchmark script that exercises its public API.

benchmark script
$ python ao_bench.py

Every reported field identical across 10 checked graphs

Optimized code & patchautooptm/networkx-ao
13.15×
RTX 4090 inferencenoise floor 1.3%

Unit timed: one McMaster image: read → add noise → SwinIR-M colour denoising → PNG written → PSNR/SSIM

command
$ python main_test_swinir.py --task color_dn --noise 15 --model_path model_zoo/swinir/005_colorDN_DFWB_s128w8_SwinIR-M_noise15.pth --folder_gt testsets/McMaster

Every output PNG within 2 levels of stock; PSNR against ground truth moves ≤0.01 dB

Optimized code & patchautooptm/SwinIR-ao
9.16×
CPU data processingnoise floor 11.4%

Unit timed: one week of a climate summary chain (coarsen → rolling → groupby_bins → anomaly → daily max → NetCDF write)
A library with no entry point: timing covers a benchmark script that exercises its public API.

benchmark script
$ python ao_bench.py

Statistics within 1.1e-5 absolute of the stock chain

Optimized code & patchautooptm/xarray-ao
8.56×
RTX 4090 inferencenoise floor 2.05%

Unit timed: one batch of 32 images through ViT-B/32 zero-shot classification (decode → preprocess → image tower → probabilities)

command
$ python main_4cf6e0.py

Zero-shot accuracy 301/1034, identical to stock

Optimized code & patchautooptm/CLIP-ao
7.48×
NVIDIA GPU inferencenoise floor 1.1%

Unit timed: one sample (200 generated tokens, pythia-410m, bf16)

command
$ python litgpt/__main__.py generate checkpoints/EleutherAI/pythia-410m --max_new_tokens 200 --num_samples 8

Bit-identical against the frozen bf16 reference, pinned and held-out prompts

Optimized code & patchautooptm/litgpt-ao

open_clip · CLAP

mlfoundations/open_clip ↗
5.82×
NVIDIA GPU inferencenoise floor 0.63%

Unit timed: one batch of 64 ESC-50 clips through CLAP zero-shot classification (audio → log-mel → audio tower → scores)

command
$ python -m open_clip_train.main --model CLAP-HTSAT-tiny-Roberta-base-fused --pretrained laion --audio-zeroshot-dataset ashraq/esc50 …

ESC-50 zero-shot top-1 87.85% / top-5 98.85%, identical to stock

Optimized code & patchautooptm/open_clip-ao

statsforecast

Nixtla/statsforecast ↗
5.24×
CPU data processingnoise floor 1.1%

Unit timed: one batch of series fitted and forecast by StatsForecast
A library with no entry point: timing covers a benchmark script that exercises its public API.

benchmark script
$ python ao_bench.py

Bit-identical forecasts

Optimized code & patchautooptm/statsforecast-ao
5.22×
RTX 4090 D inferencenoise floor 0.43%

Unit timed: one DINOv2 ViT-L/14 forward over a 16×3×518×518 batch (random weights, built from the repo's own hubconf)

command
$ python bench_dino.py

Output features within cosine 0.99998 / relative L2 0.0065 of the stock fp32 program

Optimized code & patchautooptm/dinov2-ao
4.90×
RTX 4090 inferencenoise floor 1.5%

Unit timed: one 100-frame chunk of a REDS clip: frames read, BasicVSR++ 4x video super-resolution, 100 PNGs written

command
$ python inference/inference_basicvsrpp.py

Written PNGs within 1 grey level of the stock program's on 99.9% of pixels (PSNR 61.9 dB against it).

Optimized code & patchautooptm/BasicSR-ao
4.58×
RTX 4090 inferencenoise floor 2.5%

Unit timed: one sentence synthesised end to end (phonemiser → Tacotron2-DDC → vocoder → waveform)

command
$ python TTS/bin/synthesize.py --text "The quick brown fox jumps over the lazy dog, again and again." --model_name tts_models/en/ljspeech/tacotron2-DDC --out_path out.wav --use_cuda

Bit-identical waveform (PSNR 124 dB) on pinned and held-out sentences

Optimized code & patchautooptm/coqui-ai-TTS-ao

gated_attention

qiuzh20/gated_attention ↗
3.80×
RTX 4090 inferencenoise floor 0.48%

Unit timed: one prompt: 96 new tokens greedy-generated by the 1.7B gated-attention Qwen3 checkpoint (batch 1)

command
$ python demo.py --generate --prompts prompts.txt

Next-token probabilities within 0.0078 of the stock fp32 program (relative L2 0.0028)

Optimized code & patchautooptm/gated_attention-ao
3.75×
A10 inferencenoise floor 0.8%

Unit timed: one frame (8 in-the-wild demo frames, batches of 4): person crop → Sapiens 0.3B 308-keypoint pose → keypoints written

command
$ python lite/demo/vis_pose.py sapiens_0.3b_goliath_best_goliath_AP_573_torchscript.pt2 --input <frames> --output-root <out> …

Pose heatmaps within 0.011 of the stock fp32 model (PSNR 82.8 dB), within 0.014 on held-out batch sizes.

Optimized code & patchautooptm/sapiens-ao

vision · ResNet-50

pytorch/vision ↗
3.44×
NVIDIA GPU trainingnoise floor 2.25%

Unit timed: one ResNet-50 training step at the reference recipe's batch size

command
$ python train.py --model resnet50

Verified against the frozen stock reference, including a held-out set

Optimized code & patchautooptm/vision-ao

mmaction2 · TSN

open-mmlab/mmaction2 ↗
3.41×
NVIDIA GPU trainingnoise floor 0.19%

Unit timed: one TSN R50 training iteration on Kinetics-400 at the config's batch size

command
$ python tools/train.py configs/recognition/tsn/tsn_imagenet-pretrained-r50_8xb32-1x1x3-100e_kinetics400-rgb.py

Verified against the frozen stock reference, including a held-out set

Optimized code & patchautooptm/mmaction2-ao
3.37×
H100 inferencenoise floor 0.29%

Unit timed: one demo frame: person detection → Sapiens2-1B 308-keypoint pose → skeleton overlay written

command
$ python tools/vis/vis_pose.py <detector> configs/keypoints308/…/sapiens2_1b_keypoints308_shutterstock_goliath_3po-1024x768.py <pose.safetensors> …

Pose heatmaps within 0.045 of the stock fp32 model (PSNR 85.0 dB), within 0.033 on a held-out frame; all 100 frames written.

Optimized code & patchautooptm/sapiens2-ao
3.35×
NVIDIA GPU inferencenoise floor 0.8%

Unit timed: one input image: decode → 4x RRDBNet (fp16) → post-process → encode + write

command
$ python inference_realesrgan.py -n RealESRGAN_x4plus -i inputs

PSNR 59.8 dB against the stock output, under one 8-bit code

Optimized code & patchautooptm/Real-ESRGAN-ao
3.31×
RTX 4090 inferencenoise floor 3.4%

Unit timed: one KITTI LiDAR scan through the PV-RCNN demo: read, voxelise, detect

command
$ python demo.py --cfg_file cfgs/kitti_models/pv_rcnn.yaml --ckpt pv_rcnn_8369.pth --data_path ${POINT_CLOUD_DATA}

Detections within 6e-4 of the stock program's, 170x below the model's 0.1 score threshold, also on held-out scans.

Optimized code & patchautooptm/OpenPCDet-ao
3.29×
RTX 4090 inferencenoise floor 0.30%

Unit timed: one image through PP-OCRv6 medium: decode → text detection → recognition → CTC decode

command
$ python run_ocr.py

Text and boxes identical on held-out images; the most sensitive pinned page differs by one or two characters (relative L2 0.054)

Optimized code & patchautooptm/PaddleOCR-ao
3.29×
RTX 4090 inferencenoise floor 0.32%

Unit timed: one image sequence: decode → recurrent CUT3R reconstruction → per-frame depth, confidence, colour and camera files written

command
$ python demo.py --model_path src/cut3r_512_dpt_4_64.pth --seq_path examples/001 --size 512 --vis_threshold 1.5 --output_dir tmp

Point clouds within relative L2 1.9e-4 of stock; written PNGs decode to the same pixels

Optimized code & patchautooptm/CUT3R-ao
2.68×
CPU data processingnoise floor 6.9%

Unit timed: one output frame of the 10-minute-tutorial trailer rendered end to end (decode → effects → composite → ffmpeg encode)

command
$ python docs/_static/code/getting_started/moviepy_10_minutes/trailer.py

At most 1 code of pixel difference (PSNR 51.7 dB)

Optimized code & patchautooptm/moviepy-ao

yolov5 · detect

ultralytics/yolov5 ↗
2.62×
RTX 4090 inferencenoise floor 0.9%

Unit timed: one image through detect.py: read → letterbox → yolov5s forward → NMS → annotated write

command
$ python detect.py --source data/images --weights yolov5s.pt

Bit-identical on pinned images and holdout

Optimized code & patchautooptm/yolov5-ao

shallow-vs-deep-alignment

Unispac/shallow-vs-deep-alignment ↗
2.61×
RTX 4090 trainingnoise floor 0.5%

Unit timed: one optimizer step of supervised fine-tuning of TinyLlama-1.1B-Chat on GSM8K (batch 16, bf16)

command
$ python finetune.py --model_name_or_path=ckpts/tinyllama-1.1b-chat --dataset_name=gsm8k --model_family=llama2 …

Per-step training loss within 0.11% of the stock program's (its own repeat runs differ by 0.15%); gradient cosine 0.9999.

Optimized code & patchautooptm/shallow-vs-deep-alignment-ao

transformers · GPT-2

huggingface/transformers ↗
2.26×
RTX 5090 trainingnoise floor 1.5%

Unit timed: the whole training loop: 71 optimizer steps of GPT-2 (124M) on wikitext-2, warm-up included

command
$ python run_clm.py --model_name_or_path openai-community/gpt2 --dataset_name wikitext --dataset_config_name wikitext-2-raw-v1 …

Train loss 3.387 vs 3.376 stock, eval perplexity 19.74 vs 19.55

Optimized code & patchautooptm/transformers-ao
2.25×
H100 trainingnoise floor 1.24%

Unit timed: one DiT-XL/2 training step, batch 16 on one GPU

command
$ torchrun --nnodes=1 --nproc_per_node=1 train.py --model DiT-XL/2 --data-path <train dir> --global-batch-size 16 …

Per-step loss within 1.6e-4 of stock fp32; gradients match (cosine 1.000, relative L2 0.004), on held-out batches too.

Optimized code & patchautooptm/DiT-ao
2.23×
RTX 4090 inferencenoise floor 0.3%

Unit timed: one COCO image, from decode to rescaled boxes

command
$ python ultralytics/cfg/__init__.py predict model=yolov12n.pt source=data/val2017 …

Detections bit-identical to the stock program (max abs diff 0), also on 24 held-out images.

Optimized code & patchautooptm/yolov12-ao
1.98×
RTX 4090 inferencenoise floor 0.4%

Unit timed: one generated 256×256 image (250 sampling steps + VAE decode)

command
$ python inference.py --config configs/reproductions/lightningdit_xl_vavae_f16d32_64ep_cfg.yaml --demo

Images match the stock output at 41.6 dB PSNR (40.0 dB on held-out batch sizes).

Optimized code & patchautooptm/LightningDiT-ao
1.90×
RTX 5090 inferencenoise floor 2.85%

Unit timed: one clip: decode → upload to the GPU → CoTracker3 offline tracks a 10×10 grid → tracks drawn on every frame → mp4 written

command
$ python demo.py --grid_size 10

Tracked points, visibility flags and rendered marker colours bit-identical to the stock program's, also on held-out clips and grid sizes

Optimized code & patchautooptm/co-tracker-ao
1.86×
RTX 4090 inferencenoise floor 2.1%

Unit timed: one clip scored with detect-adaptive end to end (decode → per-frame score → cut decision → scene list)

command
$ python scenedetect/__main__.py -i demo.mp4 detect-adaptive list-scenes -n

Bit-identical scores and an identical cut list

Optimized code & patchautooptm/PySceneDetect-ao
1.84×
RTX 4090 inferencenoise floor 4.0%

Unit timed: one image + prompt through demo/inference_on_a_image.py (SwinT backbone, BERT, deformable decoder → boxes)

command
$ python demo/inference_on_a_image.py -c groundingdino/config/GroundingDINO_SwinT_OGC.py -p weights/groundingdino_swint_ogc.pth -i .asset/cat_dog.jpeg -o out -t "cat ear."

Boxes move by less than 0.1 px

Optimized code & patchautooptm/GroundingDINO-ao
1.82×
RTX 4090 inferencenoise floor 0.86%

Unit timed: one DAVIS 2017 val video tracked end to end with SAM 2.1 hiera-b+: frames decoded, frame-0 masks added, every frame propagated, label PNGs written

command
$ python ./tools/vos_inference.py --sam2_cfg configs/sam2.1/sam2.1_hiera_b+.yaml --sam2_checkpoint ./checkpoints/sam2.1_hiera_base_plus.pt --base_video_dir <DAVIS>/JPEGImages/480p … --output_mask_dir ./outputs/davis_2017_pred_pngs

Written label PNGs identical to the stock program's on at least 99.9% of pixels (only a few object-boundary pixels change label), also on held-out videos

Optimized code & patchautooptm/sam2-ao
1.65×
RTX 5090 trainingnoise floor 3.4%

Unit timed: one REPA-E training step at 256×256 (SiT-B/2 + SD-VAE + DINOv2-B), batch 4

command
$ python train_repae.py --model=SiT-B/2 --vae=f8d4 --enc-type=dinov2-vit-b --batch-size=4 --mixed-precision=fp16 …

Per-step loss within 0.9% of stock, inside what the command's own fp16 already moves

Optimized code & patchautooptm/REPA-E-ao
1.63×
RTX 5090 inferencenoise floor 0.20%

Unit timed: one 1280×720 JPEG frame: decode → SAM 3 text-prompted detection + segmentation ("person") → 0.5 threshold, masks upsampled to full size

command
$ python perf/confirm_e2e.py --arm baseline --frames 64

Detection verdicts identical on every query; scores and masks within cosine 0.99999 of stock

Optimized code & patchautooptm/sam3-ao
1.58×
RTX 4090 inferencenoise floor 1.7%

Unit timed: one scene reconstruction (decode, VGGT-1B, unprojection, COLMAP)

command
$ python demo_colmap.py --scene_dir=examples/kitchen

World point cloud matches the stock program's at PSNR 97.2 dB (relative L2 1.8e-4).

Optimized code & patchautooptm/vggt-ao
1.56×
RTX 4090 inferencenoise floor 0.21%

Unit timed: one clip streamed to the faster_whisper server over a fresh websocket session (first packet → transcript covering the clip → end of audio), Whisper small
A library with no entry point: timing covers a benchmark script that exercises its public API.

benchmark script
$ python ao_bench.py

Transcripts identical to the stock program's, word for word, on pinned and held-out clips

Optimized code & patchautooptm/WhisperLive-ao
1.49×
RTX 4090 trainingnoise floor 0.05%

Unit timed: one optimizer step (8 x 64x128-token micro-batches, 65,536 tokens)

command
$ python main.py loader.global_batch_size=512 … model=small algo=bd3lm data=lm1b-wrap model.length=128 block_size=16 mode=train …

Training loss within 1.1% of the stock program's on the same batches; 0.77% on held-out batches.

Optimized code & patchautooptm/bd3lms-ao
1.47×
RTX 4090 trainingnoise floor 0.23%

Unit timed: one LJSpeech training step (generator → MPD + MSD discriminators → generator update), batch 16

command
$ python train.py --config config_v1.json

Mel L1 after two epochs 0.544 / 0.539 vs stock 0.545 / 0.544

Optimized code & patchautooptm/hifi-gan-ao
1.41×
RTX 4090 trainingnoise floor 0.2%

Unit timed: one training iteration of train.py (shakespeare_char)

command
$ python train.py config/train_shakespeare_char.py --compile=False --max_iters=400 --eval_interval=200 --log_interval=10

Loss trajectory within 8e-5 relative, gradient cosine ≥ 0.99999

Optimized code & patchautooptm/nanoGPT-ao
1.39×
RTX 4090 inferencenoise floor 1.3%

Unit timed: one photo: face detection, restoration, 2x background upsampling, write-out

command
$ python inference_gfpgan.py -i inputs/whole_imgs -o results -v 1.3 -s 2

Restored images at PSNR 48.4 dB against the stock program's; 47.4 dB on held-out photos.

Optimized code & patchautooptm/GFPGAN-ao
1.35×
RTX 5090 trainingnoise floor 0.4%

Unit timed: one LoRA fine-tuning step of VibeVoice-ASR on the toy dataset (decode → tokenise → collate → forward/backward), measured on an RTX 5090

command
$ python finetuning-asr/lora_finetune.py --model_path microsoft/VibeVoice-ASR --data_dir finetuning-asr/toy_dataset --output_dir out --num_train_epochs 3 --per_device_train_batch_size 1 --learning_rate 1e-4 --bf16 --report_to none

Loss within 6.5e-2 relative, gradient cosine 0.885; a switch restores the stock path, the user decides

Optimized code & patchautooptm/VibeVoice-ao
1.28×
RTX 5090 trainingnoise floor 0.12%

Unit timed: one pre-training step of tabicl.train (stage-1 recipe, batch 32 synthetic tables from the mix-SCM prior, bf16)

command
$ python bench_train.py

Per-step loss within 2e-5 relative of the stock program's (its own repeat-run spread); gradient cosine 1.0

Optimized code & patchautooptm/tabicl-ao

yolov5 · train

ultralytics/yolov5 ↗
1.24×
RTX 4090 trainingnoise floor 2.0%

Unit timed: one training step of train.py on a batch of 16 coco128 images

command
$ python train.py --data coco128.yaml --weights yolov5s.pt --img 640 --epochs 3 --batch-size 16

Loss within 2e-3 relative, gradient cosine ≥ 0.996

Optimized code & patchautooptm/yolov5-ao

mmdetection · Faster R-CNN

open-mmlab/mmdetection ↗
1.22×
NVIDIA GPU trainingnoise floor 1.38%

Unit timed: one Faster R-CNN R50-FPN training iteration on COCO at the config's batch size

command
$ python tools/train.py configs/faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py

Verified against the frozen stock reference, including a held-out set

Optimized code & patchautooptm/mmdetection-ao
1.22×
A100-40GB training

Unit timed: one CDiT-XL/2 bf16 training step, batch 4 at 224 px on synthetic trajectories

command
$ python train.py --config config/ao_synth.yaml --epochs 2 --bfloat16 1 --torch-compile 0 …

Loss within 1.2e-7 of stock; gradient cosine 1.0, inside the stock program's own run-to-run spread

Optimized code & patchautooptm/nwm-ao
1.19×
RTX 4090 inferencenoise floor 8.4%

Unit timed: one audio file transcribed (--model small, float16, no alignment)

command
$ python whisperx/__main__.py audio.wav --model small --compute_type float16 --no_align --output_dir out

Features within 0.016 of the stock implementation

Optimized code & patchautooptm/whisperX-ao
1.18×
RTX 4090 trainingnoise floor 2.3%

Unit timed: one series: TimeSeries → Scaler → NBEATSModel → fit (3 epochs) → forecast
A library with no entry point: timing covers a benchmark script that exercises its public API.

benchmark script
$ python ao_bench.py

Loss within 2e-3 relative, MAPE unchanged

Optimized code & patchautooptm/darts-ao

How we measure

Same command, same inputs, same output

01
Timed end to end

We time the wall-clock duration of one unit of work under the project's own command — including data loading, preprocessing and writes — not just the model's forward pass.

02
Baseline frozen first

Before any change is made, the inputs are pinned, a reference output is recorded and the host's noise floor is measured. Every change is then compared against that frozen baseline, using the median of repeated samples.

03
Outputs verified item by item

Inference projects compare the output itself (bit-identical, PSNR or cosine similarity); training projects compare the loss trajectory and gradients. Any change outside tolerance is reverted and never counted.

04
Your command stays the same

You receive a standard patch: the same files, the same flags, the same outputs. Optimizations are on by default, and each change comes with a switch that restores the original code path.

These are runs we submitted ourselves on public repositories; the hardware and command are listed on each card. The same code performs differently on other hardware or with other inputs, which is why your repository gets a free estimate first — and the decision to continue is yours.

See how much faster your repository can run

The estimate is free, takes seconds, and runs none of your code.