csvkit
wireservice/csvkit ↗Unit timed: one csvstat run over a 550k-row CSV (end to end)
$ python csvkit/utilities/csvstat.py big.csvBit-identical: every reported statistic equals stock csvkit
Optimized code & patchautooptm/csvkit-aoMeasured results
We ran AutoOptm on these public PyTorch and Python repositories exactly as we handle any submission: the project's own command (or, for a library with no entry point, a benchmark script exercising its public API), timed end to end, with outputs compared item by item before and after.
Unit timed: one csvstat run over a 550k-row CSV (end to end)
$ python csvkit/utilities/csvstat.py big.csvBit-identical: every reported statistic equals stock csvkit
Optimized code & patchautooptm/csvkit-aoUnit timed: one random graph through six analyses (betweenness, PageRank, clustering, …)
A library with no entry point: timing covers a benchmark script that exercises its public API.
$ python ao_bench.pyEvery reported field identical across 10 checked graphs
Optimized code & patchautooptm/networkx-aoUnit timed: one McMaster image: read → add noise → SwinIR-M colour denoising → PNG written → PSNR/SSIM
$ python main_test_swinir.py --task color_dn --noise 15 --model_path model_zoo/swinir/005_colorDN_DFWB_s128w8_SwinIR-M_noise15.pth --folder_gt testsets/McMasterEvery output PNG within 2 levels of stock; PSNR against ground truth moves ≤0.01 dB
Optimized code & patchautooptm/SwinIR-aoUnit timed: one week of a climate summary chain (coarsen → rolling → groupby_bins → anomaly → daily max → NetCDF write)
A library with no entry point: timing covers a benchmark script that exercises its public API.
$ python ao_bench.pyStatistics within 1.1e-5 absolute of the stock chain
Optimized code & patchautooptm/xarray-aoUnit timed: one batch of 32 images through ViT-B/32 zero-shot classification (decode → preprocess → image tower → probabilities)
$ python main_4cf6e0.pyZero-shot accuracy 301/1034, identical to stock
Optimized code & patchautooptm/CLIP-aoUnit timed: one test image: JPEG decode → transform → generator → PNG written to disk
$ python test.py --dataroot datasets/horse2zebra/testA --name horse2zebra_pretrained --model test --no_dropoutWorst pixel differs by 2 uint8 levels (PSNR 64.9 dB), less than the stock program differs from itself
Optimized code & patchautooptm/pytorch-CycleGAN-and-pix2pix-aoUnit timed: one sample (200 generated tokens, pythia-410m, bf16)
$ python litgpt/__main__.py generate checkpoints/EleutherAI/pythia-410m --max_new_tokens 200 --num_samples 8Bit-identical against the frozen bf16 reference, pinned and held-out prompts
Optimized code & patchautooptm/litgpt-aoUnit timed: one batch of 64 ESC-50 clips through CLAP zero-shot classification (audio → log-mel → audio tower → scores)
$ python -m open_clip_train.main --model CLAP-HTSAT-tiny-Roberta-base-fused --pretrained laion --audio-zeroshot-dataset ashraq/esc50 …ESC-50 zero-shot top-1 87.85% / top-5 98.85%, identical to stock
Optimized code & patchautooptm/open_clip-aoUnit timed: one batch of series fitted and forecast by StatsForecast
A library with no entry point: timing covers a benchmark script that exercises its public API.
$ python ao_bench.pyBit-identical forecasts
Optimized code & patchautooptm/statsforecast-aoUnit timed: one DINOv2 ViT-L/14 forward over a 16×3×518×518 batch (random weights, built from the repo's own hubconf)
$ python bench_dino.pyOutput features within cosine 0.99998 / relative L2 0.0065 of the stock fp32 program
Optimized code & patchautooptm/dinov2-aoUnit timed: one 100-frame chunk of a REDS clip: frames read, BasicVSR++ 4x video super-resolution, 100 PNGs written
$ python inference/inference_basicvsrpp.pyWritten PNGs within 1 grey level of the stock program's on 99.9% of pixels (PSNR 61.9 dB against it).
Optimized code & patchautooptm/BasicSR-aoUnit timed: one sentence synthesised end to end (phonemiser → Tacotron2-DDC → vocoder → waveform)
$ python TTS/bin/synthesize.py --text "The quick brown fox jumps over the lazy dog, again and again." --model_name tts_models/en/ljspeech/tacotron2-DDC --out_path out.wav --use_cudaBit-identical waveform (PSNR 124 dB) on pinned and held-out sentences
Optimized code & patchautooptm/coqui-ai-TTS-aoUnit timed: one prompt: 96 new tokens greedy-generated by the 1.7B gated-attention Qwen3 checkpoint (batch 1)
$ python demo.py --generate --prompts prompts.txtNext-token probabilities within 0.0078 of the stock fp32 program (relative L2 0.0028)
Optimized code & patchautooptm/gated_attention-aoUnit timed: one frame (8 in-the-wild demo frames, batches of 4): person crop → Sapiens 0.3B 308-keypoint pose → keypoints written
$ python lite/demo/vis_pose.py sapiens_0.3b_goliath_best_goliath_AP_573_torchscript.pt2 --input <frames> --output-root <out> …Pose heatmaps within 0.011 of the stock fp32 model (PSNR 82.8 dB), within 0.014 on held-out batch sizes.
Optimized code & patchautooptm/sapiens-aoUnit timed: one 33-frame clip encoded by the Wan 2.1 VAE (diffusers)
$ python wan21_encode.pyLatents within PSNR 73.6 dB of the fp32 encode
Optimized code & patchautooptm/ao-vae-bench-aoUnit timed: one ResNet-50 training step at the reference recipe's batch size
$ python train.py --model resnet50Verified against the frozen stock reference, including a held-out set
Optimized code & patchautooptm/vision-aoUnit timed: one training iteration on the Tanks&Temples train scene (render → loss → backward → densify → step)
$ python train.py -s tandt_db/tandt/train -m output/train --iterations 7000Per-iteration loss within 2.5e-6 relative of stock
Optimized code & patchautooptm/gaussian-splatting-aoUnit timed: one TSN R50 training iteration on Kinetics-400 at the config's batch size
$ python tools/train.py configs/recognition/tsn/tsn_imagenet-pretrained-r50_8xb32-1x1x3-100e_kinetics400-rgb.pyVerified against the frozen stock reference, including a held-out set
Optimized code & patchautooptm/mmaction2-aoUnit timed: one demo frame: person detection → Sapiens2-1B 308-keypoint pose → skeleton overlay written
$ python tools/vis/vis_pose.py <detector> configs/keypoints308/…/sapiens2_1b_keypoints308_shutterstock_goliath_3po-1024x768.py <pose.safetensors> …Pose heatmaps within 0.045 of the stock fp32 model (PSNR 85.0 dB), within 0.033 on a held-out frame; all 100 frames written.
Optimized code & patchautooptm/sapiens2-aoUnit timed: one input image: decode → 4x RRDBNet (fp16) → post-process → encode + write
$ python inference_realesrgan.py -n RealESRGAN_x4plus -i inputsPSNR 59.8 dB against the stock output, under one 8-bit code
Optimized code & patchautooptm/Real-ESRGAN-aoUnit timed: one KITTI LiDAR scan through the PV-RCNN demo: read, voxelise, detect
$ python demo.py --cfg_file cfgs/kitti_models/pv_rcnn.yaml --ckpt pv_rcnn_8369.pth --data_path ${POINT_CLOUD_DATA}Detections within 6e-4 of the stock program's, 170x below the model's 0.1 score threshold, also on held-out scans.
Optimized code & patchautooptm/OpenPCDet-aoUnit timed: one image through PP-OCRv6 medium: decode → text detection → recognition → CTC decode
$ python run_ocr.pyText and boxes identical on held-out images; the most sensitive pinned page differs by one or two characters (relative L2 0.054)
Optimized code & patchautooptm/PaddleOCR-aoUnit timed: one image sequence: decode → recurrent CUT3R reconstruction → per-frame depth, confidence, colour and camera files written
$ python demo.py --model_path src/cut3r_512_dpt_4_64.pth --seq_path examples/001 --size 512 --vis_threshold 1.5 --output_dir tmpPoint clouds within relative L2 1.9e-4 of stock; written PNGs decode to the same pixels
Optimized code & patchautooptm/CUT3R-aoUnit timed: one encode() request (a list of texts → unit-norm embeddings), as the example script issues it
$ python examples/sentence_transformer/applications/computing-embeddings/computing_embeddings.pyCosine 0.99999 against the fp32 output
Optimized code & patchautooptm/sentence-transformers-aoUnit timed: one output frame of the 10-minute-tutorial trailer rendered end to end (decode → effects → composite → ffmpeg encode)
$ python docs/_static/code/getting_started/moviepy_10_minutes/trailer.pyAt most 1 code of pixel difference (PSNR 51.7 dB)
Optimized code & patchautooptm/moviepy-aoUnit timed: one ResNet-50 training step (end to end)
$ python train.py <data> --model resnet50Verified against the frozen stock reference, including a held-out set
Optimized code & patchautooptm/pytorch-image-models-aoUnit timed: one image through detect.py: read → letterbox → yolov5s forward → NMS → annotated write
$ python detect.py --source data/images --weights yolov5s.ptBit-identical on pinned images and holdout
Optimized code & patchautooptm/yolov5-aoUnit timed: one optimizer step of supervised fine-tuning of TinyLlama-1.1B-Chat on GSM8K (batch 16, bf16)
$ python finetune.py --model_name_or_path=ckpts/tinyllama-1.1b-chat --dataset_name=gsm8k --model_family=llama2 …Per-step training loss within 0.11% of the stock program's (its own repeat runs differ by 0.15%); gradient cosine 0.9999.
Optimized code & patchautooptm/shallow-vs-deep-alignment-aoUnit timed: one batch-1 autoregressive decode run through DeepSpeed's inference engine
$ python benchmark/deepspeed/run.pyVerified against the frozen stock reference, including a held-out set
Optimized code & patchautooptm/DeepSpeed-aoUnit timed: one nerfacto training iteration on a 4096-ray batch
$ ns-train nerfacto --data data/blender/legoTraining loss within 0.011% of the stock program's per step, about 0.0005 dB of rendered PSNR.
Optimized code & patchautooptm/nerfstudio-aoUnit timed: the whole training loop: 71 optimizer steps of GPT-2 (124M) on wikitext-2, warm-up included
$ python run_clm.py --model_name_or_path openai-community/gpt2 --dataset_name wikitext --dataset_config_name wikitext-2-raw-v1 …Train loss 3.387 vs 3.376 stock, eval perplexity 19.74 vs 19.55
Optimized code & patchautooptm/transformers-aoUnit timed: one DiT-XL/2 training step, batch 16 on one GPU
$ torchrun --nnodes=1 --nproc_per_node=1 train.py --model DiT-XL/2 --data-path <train dir> --global-batch-size 16 …Per-step loss within 1.6e-4 of stock fp32; gradients match (cosine 1.000, relative L2 0.004), on held-out batches too.
Optimized code & patchautooptm/DiT-aoUnit timed: one COCO image, from decode to rescaled boxes
$ python ultralytics/cfg/__init__.py predict model=yolov12n.pt source=data/val2017 …Detections bit-identical to the stock program (max abs diff 0), also on 24 held-out images.
Optimized code & patchautooptm/yolov12-aoUnit timed: one clip encoded by the Wan 2.2 VAE (diffusers)
$ python wan22_encode.pyLatents within PSNR 78.3 dB of the fp32 encode
Optimized code & patchautooptm/ao-vae-bench-aoUnit timed: one generated 256×256 image (250 sampling steps + VAE decode)
$ python inference.py --config configs/reproductions/lightningdit_xl_vavae_f16d32_64ep_cfg.yaml --demoImages match the stock output at 41.6 dB PSNR (40.0 dB on held-out batch sizes).
Optimized code & patchautooptm/LightningDiT-aoUnit timed: one clip: decode → upload to the GPU → CoTracker3 offline tracks a 10×10 grid → tracks drawn on every frame → mp4 written
$ python demo.py --grid_size 10Tracked points, visibility flags and rendered marker colours bit-identical to the stock program's, also on held-out clips and grid sizes
Optimized code & patchautooptm/co-tracker-aoUnit timed: one clip scored with detect-adaptive end to end (decode → per-frame score → cut decision → scene list)
$ python scenedetect/__main__.py -i demo.mp4 detect-adaptive list-scenes -nBit-identical scores and an identical cut list
Optimized code & patchautooptm/PySceneDetect-aoUnit timed: one image + prompt through demo/inference_on_a_image.py (SwinT backbone, BERT, deformable decoder → boxes)
$ python demo/inference_on_a_image.py -c groundingdino/config/GroundingDINO_SwinT_OGC.py -p weights/groundingdino_swint_ogc.pth -i .asset/cat_dog.jpeg -o out -t "cat ear."Boxes move by less than 0.1 px
Optimized code & patchautooptm/GroundingDINO-aoUnit timed: one DAVIS 2017 val video tracked end to end with SAM 2.1 hiera-b+: frames decoded, frame-0 masks added, every frame propagated, label PNGs written
$ python ./tools/vos_inference.py --sam2_cfg configs/sam2.1/sam2.1_hiera_b+.yaml --sam2_checkpoint ./checkpoints/sam2.1_hiera_base_plus.pt --base_video_dir <DAVIS>/JPEGImages/480p … --output_mask_dir ./outputs/davis_2017_pred_pngsWritten label PNGs identical to the stock program's on at least 99.9% of pixels (only a few object-boundary pixels change label), also on held-out videos
Optimized code & patchautooptm/sam2-aoUnit timed: one txt2img request (SD 1.5, 20 steps, 512-768 px) to PNG response
$ python launch.py --api --nowebuiDecoded images at PSNR 48.2 dB against the stock server's; 51.6 dB on held-out requests.
Optimized code & patchautooptm/stable-diffusion-webui-aoUnit timed: one REPA-E training step at 256×256 (SiT-B/2 + SD-VAE + DINOv2-B), batch 4
$ python train_repae.py --model=SiT-B/2 --vae=f8d4 --enc-type=dinov2-vit-b --batch-size=4 --mixed-precision=fp16 …Per-step loss within 0.9% of stock, inside what the command's own fp16 already moves
Optimized code & patchautooptm/REPA-E-aoUnit timed: one 33-frame clip encoded by the HunyuanVideo VAE (diffusers)
$ python hunyuan_encode.pyLatents within PSNR 50.9 dB of the fp32 encode
Optimized code & patchautooptm/ao-vae-bench-aoUnit timed: one short clip interpolated 2x end to end: decode → SSIM check → IFNet → encode
$ python inference_video.py --video=demo.mp4 --exp=1Worst pixel 1 code off (PSNR 57.7 dB)
Optimized code & patchautooptm/ECCV2022-RIFE-aoUnit timed: one 1280×720 JPEG frame: decode → SAM 3 text-prompted detection + segmentation ("person") → 0.5 threshold, masks upsampled to full size
$ python perf/confirm_e2e.py --arm baseline --frames 64Detection verdicts identical on every query; scores and masks within cosine 0.99999 of stock
Optimized code & patchautooptm/sam3-aoUnit timed: one scene reconstruction (decode, VGGT-1B, unprojection, COLMAP)
$ python demo_colmap.py --scene_dir=examples/kitchenWorld point cloud matches the stock program's at PSNR 97.2 dB (relative L2 1.8e-4).
Optimized code & patchautooptm/vggt-aoUnit timed: one clip streamed to the faster_whisper server over a fresh websocket session (first packet → transcript covering the clip → end of audio), Whisper small
A library with no entry point: timing covers a benchmark script that exercises its public API.
$ python ao_bench.pyTranscripts identical to the stock program's, word for word, on pinned and held-out clips
Optimized code & patchautooptm/WhisperLive-aoUnit timed: one multi-view Fast3R ViT-L/512 reconstruction of 20 demo-video frames
$ python bench_infer.pyPassed the output check against the stock program, frozen before optimization began
Optimized code & patchautooptm/fast3r-aoUnit timed: one optimizer step (8 x 64x128-token micro-batches, 65,536 tokens)
$ python main.py loader.global_batch_size=512 … model=small algo=bd3lm data=lm1b-wrap model.length=128 block_size=16 mode=train …Training loss within 1.1% of the stock program's on the same batches; 0.77% on held-out batches.
Optimized code & patchautooptm/bd3lms-aoUnit timed: one LJSpeech training step (generator → MPD + MSD discriminators → generator update), batch 16
$ python train.py --config config_v1.jsonMel L1 after two epochs 0.544 / 0.539 vs stock 0.545 / 0.544
Optimized code & patchautooptm/hifi-gan-aoUnit timed: one training iteration of train.py (shakespeare_char)
$ python train.py config/train_shakespeare_char.py --compile=False --max_iters=400 --eval_interval=200 --log_interval=10Loss trajectory within 8e-5 relative, gradient cosine ≥ 0.99999
Optimized code & patchautooptm/nanoGPT-aoUnit timed: one clip through the VAE round trip (encode → decode → frames, diffusers)
$ python decode.pyPSNR 62.9 dB against the fp32 reference
Optimized code & patchautooptm/h3-vae-bench-aoUnit timed: one photo: face detection, restoration, 2x background upsampling, write-out
$ python inference_gfpgan.py -i inputs/whole_imgs -o results -v 1.3 -s 2Restored images at PSNR 48.4 dB against the stock program's; 47.4 dB on held-out photos.
Optimized code & patchautooptm/GFPGAN-aoUnit timed: one LoRA fine-tuning step of VibeVoice-ASR on the toy dataset (decode → tokenise → collate → forward/backward), measured on an RTX 5090
$ python finetuning-asr/lora_finetune.py --model_path microsoft/VibeVoice-ASR --data_dir finetuning-asr/toy_dataset --output_dir out --num_train_epochs 3 --per_device_train_batch_size 1 --learning_rate 1e-4 --bf16 --report_to noneLoss within 6.5e-2 relative, gradient cosine 0.885; a switch restores the stock path, the user decides
Optimized code & patchautooptm/VibeVoice-aoUnit timed: one pre-training step of tabicl.train (stage-1 recipe, batch 32 synthetic tables from the mix-SCM prior, bf16)
$ python bench_train.pyPer-step loss within 2e-5 relative of the stock program's (its own repeat-run spread); gradient cosine 1.0
Optimized code & patchautooptm/tabicl-aoUnit timed: one training step of train.py on a batch of 16 coco128 images
$ python train.py --data coco128.yaml --weights yolov5s.pt --img 640 --epochs 3 --batch-size 16Loss within 2e-3 relative, gradient cosine ≥ 0.996
Optimized code & patchautooptm/yolov5-aoUnit timed: one Faster R-CNN R50-FPN training iteration on COCO at the config's batch size
$ python tools/train.py configs/faster_rcnn/faster-rcnn_r50_fpn_1x_coco.pyVerified against the frozen stock reference, including a held-out set
Optimized code & patchautooptm/mmdetection-aoUnit timed: one CDiT-XL/2 bf16 training step, batch 4 at 224 px on synthetic trajectories
$ python train.py --config config/ao_synth.yaml --epochs 2 --bfloat16 1 --torch-compile 0 …Loss within 1.2e-7 of stock; gradient cosine 1.0, inside the stock program's own run-to-run spread
Optimized code & patchautooptm/nwm-aoUnit timed: one audio file transcribed (--model small, float16, no alignment)
$ python whisperx/__main__.py audio.wav --model small --compute_type float16 --no_align --output_dir outFeatures within 0.016 of the stock implementation
Optimized code & patchautooptm/whisperX-aoUnit timed: one series: TimeSeries → Scaler → NBEATSModel → fit (3 epochs) → forecast
A library with no entry point: timing covers a benchmark script that exercises its public API.
$ python ao_bench.pyLoss within 2e-3 relative, MAPE unchanged
Optimized code & patchautooptm/darts-aoHow we measure
We time the wall-clock duration of one unit of work under the project's own command — including data loading, preprocessing and writes — not just the model's forward pass.
Before any change is made, the inputs are pinned, a reference output is recorded and the host's noise floor is measured. Every change is then compared against that frozen baseline, using the median of repeated samples.
Inference projects compare the output itself (bit-identical, PSNR or cosine similarity); training projects compare the loss trajectory and gradients. Any change outside tolerance is reverted and never counted.
You receive a standard patch: the same files, the same flags, the same outputs. Optimizations are on by default, and each change comes with a switch that restores the original code path.
These are runs we submitted ourselves on public repositories; the hardware and command are listed on each card. The same code performs differently on other hardware or with other inputs, which is why your repository gets a free estimate first — and the decision to continue is yours.
The estimate is free, takes seconds, and runs none of your code.