GPU Requirements by Service
The sizings on this page are for LILT 6.1.0. Every LILT feature that uses a GPU is backed by one or more of the services below; a feature runs only when each of its services can be placed on your GPU nodes.- Whole GPUs, no sharing. Each service takes its GPUs entirely; leftover VRAM on a card cannot host another service.
- VRAM adds up only within one node. A 45 GB requirement is met by one 48 GB card or by two 24 GB cards in the same server, never by cards on different servers. One large card can replace several small ones.
- Compute capability is a gate, not a size. T4 (7.5) clears the batch worker and ASR but not rayman, however much VRAM the node has. Reference values: T4 7.5, A100 8.0, A10/A6000 8.6, L4/L40S 8.9, H100 9.0.
- Ampere and Gemma. Ampere cards (8.0 to 8.6: A100, A10, A6000) run Gemma at reduced throughput but are not enabled by the default installer; ask LILT before sizing Gemma on Ampere.
- One large card for Gemma. On-prem installs default to two 24 GB cards for Gemma. A single card of 48 GB or more is supported by applying the
gemma-single-largeGPU profile shipped with the installer; EKS installs default to the single-card sizing.
Recommended Production Configurations (AWS EKS)
The following are reference sizings for AWS EKS production deployments, based on configurations validated with LILT EKS customers. Three tiers are offered — Standard ★, Medium, and Large — sized to the number of concurrent users. Unless noted, a component matches the Standard tier. For a detailed per-resource cost breakdown, see the AWS EKS section below.Prod environment — three tier options | |||
Prod component | Standard ★ | Medium | Large |
Concurrent users supported | 30–50 | 100–150 | 150+ |
Database node | RDS: db.m5d.xlarge | Same as Standard | Same as Standard |
Translate / Uzbek-Chechen (rayman v4.0) / OCR + AI Review / LILT Create (Gemma) | 1× g6e.12xlarge (4× L40S, 48 vCPU, 384 GB) | Same as Standard | 1× g6e.24xlarge (4× L40S, 96 vCPU, 768 GB) |
Batch / Whisper | 1× g6.24xlarge (4× L4, 96 vCPU, 384 GB) | Same as Standard | Same as Standard |
Worker Node | none | 1× m5.8xlarge (32 vCPU, 128 GB) | none |
Total instances | 2 | 3 | 2 |
Cloud Provider Cost Estimates
AWS Commercial
Last Revised: 23 Jul 2026
AWS Commercial | ||||||||
Platform | Resource Identifier | Specifications | Quantity | Unit | Unit cost [1] | Cost [2] per month | Cost [3] per month | Notes |
AWS | m5.xlarge | 4vCPU / 16G RAM / 200GB | 1 | Hour | $0.192 | $118.93 | $104.33 | K8s Master |
AWS | r5.16xlarge | 64vCPU / 512G RAM / 1TB | 2 | Hour | $8.064 | $4,584.66 | $4,028.40 | K8s Nodes |
AWS | g6.12xlarge | 4 L4 GPUs [4] / 48vCPU / 192G RAM / 2TB / 96GiB VRAM | 1 | Hour | $4.60 | $2,893.52 | $2,346.82 | Standard Services (Translate, Batch) |
Total Cost | All v4 lang-pairs with standard service | $12.86 | $7,597.11 | $6,479.55 | ||||
AWS | g6e.4xlarge (Optional) | 1 L40S GPU / 16vCPU / 128G RAM / 2TB / 48 GiB VRAM | 1 | Hour | $3.00 | $1,817.98 | $1,541.65 | OCR + AI Review / LILT Create (Gemma) |
AWS | g4dn.2xlarge (Optional) | 1 T4 GPU / 8vCPU / 32G RAM / 1TB / 16 GiB VRAM | 1 | Hour | $0.752 | $477.85 | $426.02 | ASR (Whisper) |
Total Cost (incl. optional services) | Standard service + OCR/AI Review/Create & ASR | $16.61 | $9,892.94 | $8,447.22 | ||||
AWS Top Secret
Last Revised: 9 Feb 2026
AWS Top Secret | |||||||
Platform | Resource Identifier | Specifications | Quantity | Unit | Unit cost [1] | Cost [2] per month | Notes |
AWS | m5.xlarge | 4vCPU / 16G RAM / 200GB | 1 | Hour | $0.45 | $260.90 | K8s Master |
AWS | r5.16xlarge | 64vCPU / 512G RAM / 1TB | 2 | Hour | $15.42 | $8,199.32 | K8s Nodes |
AWS | g6.12xlarge | 4 L4 GPUs [4] / 48vCPU / 192G RAM / 2TB / 96 GiB VRAM | 1 | Hour | $9.24 | $5,694.56 | Standard Services (Translate, Batch) |
Total Cost | All v4 lang-pairs with standard service | $25.11 | $14,154.78 | ||||
AWS | g6e.4xlarge (Optional) | 1 L40S GPU / 16vCPU / 128G RAM / 2TB / 48 GiB VRAM | 1 | Hour | $6.03 | $3,532.53 | OCR + AI Review / LILT Create (Gemma) |
AWS | g4dn.2xlarge (Optional) | 1 T4 GPU / 8vCPU / 32G RAM / 1TB / 16 GiB VRAM | 1 | Hour | $1.62 | $958.48 | ASR (Whisper) |
Total Cost (incl. optional services) | Standard service + OCR/AI Review/Create & ASR | $32.76 | $18,645.79 | ||||
AWS EKS
Last Revised: 23 Jul 2026
AWS EKS | |||||||
Platform | Resource Identifier | Specifications | Quantity | Unit | Unit cost [1] | Cost [3] per month | Notes |
AWS | Amazon EKS | N/A | 1 | Hour | $0.10 | $73.00 | EKS Control Plane |
AWS | RDS db.m5d.xlarge | 4vCPU / 16G RAM | 1 | Hour | $0.419 | $305.87 [5] | Database node (RDS) |
AWS | g6e.12xlarge | 4 L40S GPUs / 48vCPU / 384G RAM / 2TB / 192 GiB VRAM | 1 | Hour | $10.49 | $5,384.70 | Translate / Uzbek-Chechen (rayman v4.0) / OCR + AI Review / LILT Create (Gemma) |
AWS | g6.24xlarge | 4 L4 GPUs / 96vCPU / 384G RAM / 2TB / 96 GiB VRAM | 1 | Hour | $6.68 | $3,404.58 | Batch / Whisper (ASR) |
Total Cost | All v4 lang-pairs with standard service | $17.69 | $9,168.15 | ||||
- Commercial: AWS Commercial Estimate
- Top Secret: AWS Top Secret Estimate
- EKS: AWS EKS Estimate
This EKS table follows the Standard tier of the Recommended Production Configurations table above: the platform runs on two GPU nodes (
g6e.12xlarge for Translate, Uzbek-Chechen rayman v4.0 and OCR; g6.24xlarge for Batch and Whisper/ASR), backed by an RDS database node — no separate CPU worker node is required at this tier.- AI Review & LILT Create — Served by the same Gemma deployment as OCR on the
g6e.12xlargenode; no additional instance is required. - OCR (Gemma) and ASR (Whisper) — In this configuration these are co-located on the standard GPU nodes above and do not require additional instances. ASR with the Sorani (Kurdish) model needs 25 GB of VRAM on one node and has no CPU fallback.
- Uzbek and Chechen Language Support — Provided through the neural v4.0 pipeline in rayman (which replaces the former dedicated Emma model), co-located on the
g6e.12xlargenode.
GCP
Last Revised: 23 Jul 2026
GCP | ||||||||
Platform | Resource Identifier | Specifications | Quantity | Unit | Unit cost [1] | Cost [1] per month | Cost [2] per month | Notes |
GCP | n1-standard-4 | 4vCPU / 15G RAM / 500GB | 1 | Hour | $0.19 | $138.70 | $87.38 | K8s Master |
GCP | n2-highmem-64 | 64vCPU / 512G RAM / 1TB | 2 | Hour | $8.38 | $6,120.90 | $3,856.01 | K8s Nodes |
GCP | g2-standard-48 | 4 L4 GPUs [4] / 48vCPU / 192G RAM / 2TB / 96 GiB VRAM | 1 | Hour | $4.00 | $2,921.24 | $2,445.43 | Standard Services (Translate, Batch) |
Total Cost | All v4 lang-pairs with standard service | $12.57 | $9,180.84 | $6,388.82 | ||||
GCP | g2-standard-24 (Optional) | 2 L4 GPU / 24vCPU / 96G RAM / 2TB / 48 GiB VRAM | 1 | Hour | $2.00 | $1,460.58 | $1,222.75 | OCR + AI Review / LILT Create (Gemma) |
Total Cost (incl. optional services) | Standard service + OCR/AI Review/Create | $14.57 | $10,641.42 | $7,611.57 | ||||
In addition to the standard deployment, several optional services require extra GPU resources:
- OCR, AI Review & LILT Create — All three are served by one Gemma (vLLM) deployment, which needs 45 GB of VRAM on a single node (one 48 GB card, or two 24 GB cards in the same server) and 80 GB of node RAM, as shown in the optional row above. Enabling AI Review or LILT Create on top of OCR adds no GPU.
- Uzbek and Chechen Language Support — Provided through the neural v4.0 pipeline in rayman (which replaces the former dedicated Emma model). When enabled, it requires one additional L4 or A10 GPU with at least 24GB VRAM; when disabled, rayman runs on a single GPU.
- ASR (Automatic Speech Recognition) — GPU-accelerated ASR requires one additional T4 GPU with 16 GB VRAM. By default, it runs on CPU. ASR with the Sorani (Kurdish) model needs 25 GB of VRAM on one node (two 16 GB or two 24 GB cards in the same server, or one 48 GB card) and has no CPU fallback.
Microsoft Azure
Last Revised: 23 Jul 2026
Microsoft Azure | ||||||||
|---|---|---|---|---|---|---|---|---|
Platform | Resource Identifier | Specifications | Quantity | Unit | Unit cost [1] | Cost [1] per month | Cost [2] per month | Notes |
Azure | D3 v2 | 4vCPU / 14G RAM / 512GiB | 1 | Hour | $0.29 | $213.89 | $147.58 | K8s Master |
Azure | E64as v5 | 64vCPU / 512G RAM / 512GiB | 2 | Hour | $7.23 | $5,279.36 | $3,579.93 | K8s Nodes |
Azure | Standard_NV72ads_A10_v5 | 2 A10 GPU / 72vCPU / 880G RAM / 2880GiB/ 48GB VRAM | 2 | Hour | $13.04 | $9,519.20 | $7,923.78 | Standard Services (Translate, Batch) |
Total Cost | All v4 lang-pairs with standard service | $20.56 | $15,012.45 | $11,651.29 | ||||
Azure | Standard_NV72ads_A10_v5 (Optional) | 2 A10 GPU / 72vCPU / 880G RAM / 2880GiB/ 48GB VRAM | 1 | Hour | $6.52 | $4,759.60 | $3,961.89 | OCR + AI Review / LILT Create (Gemma) |
Total Cost (incl. optional services) | Standard service + OCR/AI Review/Create | $27.08 | $19,772.05 | $15,613.18 | ||||
In addition to the standard deployment, several optional services require extra GPU resources:
- OCR, AI Review & LILT Create — All three are served by one Gemma (vLLM) deployment, which needs 45 GB of VRAM on a single node (one 48 GB card, or two 24 GB cards in the same server) and 80 GB of node RAM, as shown in the optional row above. Enabling AI Review or LILT Create on top of OCR adds no GPU. The A10 is an Ampere card (compute capability 8.6): it runs Gemma at reduced throughput and is not enabled by the default installer, so confirm with LILT before sizing on the Standard_NV72ads_A10_v5 row above, or choose an Ada-class SKU.
- Uzbek and Chechen Language Support — Provided through the neural v4.0 pipeline in rayman (which replaces the former dedicated Emma model). When enabled, it requires one additional L4 or A10 GPU with at least 24GB VRAM; when disabled, rayman runs on a single GPU.
- ASR (Automatic Speech Recognition) — GPU-accelerated ASR requires one additional T4 GPU with 16 GB VRAM. By default, it runs on CPU. ASR with the Sorani (Kurdish) model needs 25 GB of VRAM on one node (two 16 GB or two 24 GB cards in the same server, or one 48 GB card) and has no CPU fallback.

