What it does
- Audits hardware, accelerators, RAM/VRAM, model inventory, and approved storage volumes.
- Selects the right engine for the hardware and model format instead of defaulting to one stack.
- Builds a model-fit estimate before downloading or converting weights.
- Defines cache location, batching, KV cache, concurrency, memory, and endpoint settings.
- Verifies with smoke benchmarks and OpenAI-compatible API checks.