Compute Resources
Selecting the right compute tier keeps latency low while controlling costs. This guide explains the available options and strategies for different workloads.
Available Tiers
Refer to the targon.Resources reference for the full list of identifiers. They fall into:
- CPU tiers (
cpu-small→cpu-xlarge) for lightweight services, background jobs, and orchestration. - GPU tiers by accelerator family — H200, H100, B200, B300 (each
small→xlarge), plus RTX6000B (small) and RTX4090 (small→large). - VM tiers (
type=vmin inventory) for confidential GPU virtual machines. See the Virtual Machines guide. - Bare Metal hardware classes (
type=bmin inventory, scoped byregion) for dedicated physical servers. See the Bare Metal guide.
Check live availability and pricing:
targon inventory --gpu
targon inventory --type vm --gpu
Matching Workloads to Tiers
- API backends / web hooks: Start with
cpu-smallorcpu-medium. Increase tiers only if you see sustained CPU saturation. - Batch jobs / ETL: Use
cpu-largeorcpu-xlargefor parallel processing. - LLM inference: Pick a tier that matches model size and throughput — start with
h200-smalland scale up (h200-large/h200-xlarge) as needed. - Interactive development: Start with a Sandbox for a fast, isolated environment with terminal and desktop access.
Cost Optimization Tips
- Right-size workloads instead of over-provisioning GPU memory you do not use.
- Pause or delete sandboxes when they are not in use.
- Use lifecycle timeouts to clean up temporary environments automatically.
Related Reading
- Compute API reference
- Sandboxes guide for development and agent environments
- Virtual Machines guide for confidential GPU VMs
- Bare Metal guide for dedicated physical servers