GPU-accelerated bit classification from semiconductor die photographs. Automatic grid detection, fusion across every capture, one-click review of uncertain bits, and integrated decode — in a single native application.
From raw SEM image to decoded firmware. Every step handled in a single tool.
CUDA-powered bit classification via PyTorch. Processes millions of sample points in seconds. Transparent CPU fallback.
FFT-based period detection with sub-pixel refinement. No manual line placement. Handles perspective-distorted dies.
Step through uncertain bits row by row or least certain first, compare every capture side by side, and decide with one click on the reticle: 0, 1, confirm, or capture majority. You set the review threshold.
Stack multiple SEM photos of the same block. Captures are aligned automatically; confidence-weighted merging down-weights blurred or misaligned ones.
Row/column balance, columns with unusually many uncertain bits, global skew. Findings are marked on the image; F8 steps through them.
Every transform (invert, rotate, mirror) across 26 bit-order decoders. Identical outputs are merged; results ranked by entropy and known patterns.
Find reset vectors, file headers and strings across decodings. ARM, Thumb and ARM64 disassembly — click an instruction to see its bits on the die.
Diff an extraction against a known dump or another block. Every mismatching bit is listed and marked on the image.
Process all ROM blocks in a project directory. Automatic grid detection and bit extraction across the entire die.
Every bit toggle, deletion, and grid edit is tracked. Full history navigation with Ctrl+Z / Ctrl+Shift+Z.
Non-uniform column groups, horizontal row dividers, and metal trace filtering. Beyond simple uniform grids.
Export as .bin (with decoder selection), .npy, ASCII, three-class 0/1/unknown, or full-resolution composite images with bits and findings.
Real captures of a 512 × 512 mask ROM block, straight from the application.
A guided stage-by-stage workflow from raw SEM image to decoded binary.
Click four corners of the ROM area on one capture. The other captures are aligned to it.
FFT-based detection finds every row, column and divider with sub-pixel refinement.
GPU classification of every bit, fused across all captures with per-bit confidence.
Step through uncertain bits row by row and decide on the reticle while comparing every capture — or accept clear capture majorities in one step.
Design rule checks, decode ranking, pattern search, disassembly and export.
Feature comparison against existing open-source tools.
| Feature | RomScope | Bitract | Rompar | MaskRomTool |
|---|---|---|---|---|
| Grid Detection | ||||
| Automatic grid | ✓ FFT + sub-pixel | ✕ Manual polygon | ✕ Manual clicks | ✕ Manual lines |
| Perspective correction | ✓ 4-corner quad | ✕ | ✕ | ~ Tilt aligner |
| Variable-gap columns | ✓ | ~ Pattern arrays | ✕ | ✕ |
| Row divider detection | ✓ Auto + verify | ✕ | ✕ | ✕ |
| Bit Extraction | ||||
| GPU acceleration | ✓ CUDA / PyTorch | ✕ | ✕ | ✕ OpenGL view only |
| Confidence scoring | ✓ Per-bit | ~ Threshold warnings | ✕ | ~ Ambiguous flag |
| Multi-image fusion | ✓ Confidence-weighted | ✕ | ✕ | ~ ASCII diff |
| Flat-field correction | ✓ | ~ Multipliers | ✕ | ✕ |
| Analysis & Decode | ||||
| Decode configurations | ✓ 832, auto-ranked | ✕ | ✕ | ~ 5 via GatoROM |
| Interleaved decoders | ✓ 26 algorithms | ✕ | ✕ | ~ 5 algorithms |
| String / keyword search | ✓ 100+ keywords + byte patterns | ✕ | ✓ Hex search | ✕ |
| ARM disassembly | ✓ ARM / Thumb / ARM64 | ✕ | ✕ | ~ External |
| Hex view linked to image bits | ✓ | ✕ | ✓ | ✕ |
| Platform | ||||
| License | Free eval / edu + Commercial | BSD-2 | GPL-2 | GPL |
| Cross-platform | ✓ Win / Mac / Linux | ~ Windows | ✓ | ✓ |
| Actively maintained | ✓ 2026 | ✕ 2019 | ✕ 2020 | ✓ |
RomScope works on CPU-only systems but delivers peak performance with CUDA-capable hardware.
| OS | Windows 10/11, Ubuntu 22.04+, or macOS 13+ |
| CPU | 8+ cores / 16 threads (AMD Zen 3 / Intel 12th+) |
| RAM | 16–32 GB |
| GPU | NVIDIA with 4+ GB VRAM and CUDA 12.0+ (e.g. RTX 3060 or better) |
| Storage | SSD recommended for large image sets |
| Display | 1920 × 1080 or higher |
| Fallback | CPU-only mode works without GPU — slower but fully functional |
| System | 24-core x86_64, 16 GB VRAM, CUDA 13.0 |
| FFT grid detect | 12.9 ms per 10K×9K image (90 MP) |
| GPU convolution | 68 ms — 1,315 MP/s (10K×9K) |
| Bit classify | 0.88 ms for 131k bits — 150 Mb/s |
| Decode (832 cfg) | 182 ms — 4,575 configs/s (128×128 block) |
| Bit packing | 0.008 ms for 512×512 |
| Peak VRAM | 721 MB on 90 MP image |
romscope --benchmark to measure your system.
Tests include FFT grid detection, GPU/CPU convolution, bit classification (131k bits at ~150 Mb/s),
the full 832-configuration brute-force decode, and decode ranking of a 512×512 block.
Full evaluation access with no registration and no time limit. Upgrade for export, analysis, and batch automation.
Download the free evaluation. No registration, no time limit. Need commercial features? Request a license.