PremoveITN instance resident and reuse it across requests.
Separate the lifecycle costs
The benchmark excludes download and initialization. A one-shot CLI measurement
is not comparable to retained warm latency.
Select a device
device="auto" selects CUDA when available, then Apple MPS, then CPU. Pass
cpu, mps, or cuda to select a device explicitly. An unavailable explicit
device fails with a runtime error.
Release wheels are validated on macOS 14+ arm64 and manylinux_2_28 x86_64 for
Python 3.11–3.13. Frozen-model inference is validated on Apple Silicon MPS and
Linux x86-64 CPU. CUDA, Windows, macOS Intel, Linux ARM64, and other
accelerators are not validated v0.1.0 support claims.
See the platform support matrix for the exact
release boundary.
Pin production deployments
Pin the package release:PremoveITN.from_pretrained().