Skip to main content
LILT can be installed on custom hardware in a private data center or on major cloud platforms. The tables below provide approximate monthly cost estimates for running the base LILT system V4 models in a cloud environment. These estimates include long-term pricing discounts where applicable. However, resources (especially storage) may need to be increased to suit your needs. These costs reflect the cloud platform vendor’s pricing at the time of publication and are not guaranteed. Network, bandwidth, security groups, permission management, and other platform costs are not included in these estimates. Interested in installing LILT’s self-managed solution? Contact Sales from the LILT homepage.
Plan for two environments. A standard self-managed installation includes two separate environments — one test and one production. The test environment lets you validate upgrades, configuration changes, and integrations before promoting them to production. Because it mirrors the production footprint, budget for roughly double the total cloud costs shown in the estimates below, which reflect a single environment.

GPU Requirements by Service

The sizings on this page are for LILT 6.1.0. Every LILT feature that uses a GPU is backed by one or more of the services below; a feature runs only when each of its services can be placed on your GPU nodes. Placement rules:
  • Whole GPUs, no sharing. Each service takes its GPUs entirely; leftover VRAM on a card cannot host another service.
  • VRAM adds up only within one node. A 45 GB requirement is met by one 48 GB card or by two 24 GB cards in the same server, never by cards on different servers. One large card can replace several small ones.
  • Compute capability is a gate, not a size. T4 (7.5) clears the batch worker and ASR but not rayman, however much VRAM the node has. Reference values: T4 7.5, A100 8.0, A10/A6000 8.6, L4/L40S 8.9, H100 9.0.
  • Ampere and Gemma. Ampere cards (8.0 to 8.6: A100, A10, A6000) run Gemma at reduced throughput but are not enabled by the default installer; ask LILT before sizing Gemma on Ampere.
  • One large card for Gemma. On-prem installs default to two 24 GB cards for Gemma. A single card of 48 GB or more is supported by applying the gemma-single-large GPU profile shipped with the installer; EKS installs default to the single-card sizing.
The following are reference sizings for AWS EKS production deployments, based on configurations validated with LILT EKS customers. Three tiers are offered — Standard ★, Medium, and Large — sized to the number of concurrent users. Unless noted, a component matches the Standard tier. For a detailed per-resource cost breakdown, see the AWS EKS section below.

Prod environment — three tier options

Prod component

Standard ★

Medium

Large

Concurrent users supported

30–50

100–150

150+

Database node

RDS: db.m5d.xlarge

Same as Standard

Same as Standard

Translate / Uzbek-Chechen (rayman v4.0) / OCR + AI Review / LILT Create (Gemma)

1× g6e.12xlarge (4× L40S, 48 vCPU, 384 GB)

Same as Standard

1× g6e.24xlarge (4× L40S, 96 vCPU, 768 GB)

Batch / Whisper

1× g6.24xlarge (4× L4, 96 vCPU, 384 GB)

Same as Standard

Same as Standard

Worker Node

none

1× m5.8xlarge (32 vCPU, 128 GB)

none

Total instances

2

3

2

★ The Standard tier is the recommended default starting point for most deployments. As your usage grows, you can scale up to the Medium and Large tiers as needed.

Cloud Provider Cost Estimates

AWS Commercial

Last Revised: 23 Jul 2026

AWS Commercial

Platform

Resource Identifier

Specifications

Quantity

Unit

Unit cost [1]

Cost [2] per month

Cost [3] per month

Notes

AWS

m5.xlarge

4vCPU / 16G RAM / 200GB

1

Hour

$0.192

$118.93

$104.33

K8s Master

AWS

r5.16xlarge

64vCPU / 512G RAM / 1TB

2

Hour

$8.064

$4,584.66

$4,028.40

K8s Nodes

AWS

g6.12xlarge

4 L4 GPUs [4] / 48vCPU / 192G RAM / 2TB / 96GiB VRAM

1

Hour

$4.60

$2,893.52

$2,346.82

Standard Services (Translate, Batch)

Total Cost

All v4 lang-pairs with standard service

$12.86

$7,597.11

$6,479.55

AWS

g6e.4xlarge (Optional)

1 L40S GPU / 16vCPU / 128G RAM / 2TB / 48 GiB VRAM

1

Hour

$3.00

$1,817.98

$1,541.65

OCR + AI Review / LILT Create (Gemma)

AWS

g4dn.2xlarge (Optional)

1 T4 GPU / 8vCPU / 32G RAM / 1TB / 16 GiB VRAM

1

Hour

$0.752

$477.85

$426.02

ASR (Whisper)

Total Cost (incl. optional services)

Standard service + OCR/AI Review/Create & ASR

$16.61

$9,892.94

$8,447.22

AWS Top Secret

Last Revised: 9 Feb 2026

AWS Top Secret

Platform

Resource Identifier

Specifications

Quantity

Unit

Unit cost [1]

Cost [2] per month

Notes

AWS

m5.xlarge

4vCPU / 16G RAM / 200GB

1

Hour

$0.45

$260.90

K8s Master

AWS

r5.16xlarge

64vCPU / 512G RAM / 1TB

2

Hour

$15.42

$8,199.32

K8s Nodes

AWS

g6.12xlarge

4 L4 GPUs [4] / 48vCPU / 192G RAM / 2TB / 96 GiB VRAM

1

Hour

$9.24

$5,694.56

Standard Services (Translate, Batch)

Total Cost

All v4 lang-pairs with standard service

$25.11

$14,154.78

AWS

g6e.4xlarge (Optional)

1 L40S GPU / 16vCPU / 128G RAM / 2TB / 48 GiB VRAM

1

Hour

$6.03

$3,532.53

OCR + AI Review / LILT Create (Gemma)

AWS

g4dn.2xlarge (Optional)

1 T4 GPU / 8vCPU / 32G RAM / 1TB / 16 GiB VRAM

1

Hour

$1.62

$958.48

ASR (Whisper)

Total Cost (incl. optional services)

Standard service + OCR/AI Review/Create & ASR

$32.76

$18,645.79

AWS EKS

Last Revised: 23 Jul 2026

AWS EKS

Platform

Resource Identifier

Specifications

Quantity

Unit

Unit cost [1]

Cost [3] per month

Notes

AWS

Amazon EKS

N/A

1

Hour

$0.10

$73.00

EKS Control Plane

AWS

RDS db.m5d.xlarge

4vCPU / 16G RAM

1

Hour

$0.419

$305.87 [5]

Database node (RDS)

AWS

g6e.12xlarge

4 L40S GPUs / 48vCPU / 384G RAM / 2TB / 192 GiB VRAM

1

Hour

$10.49

$5,384.70

Translate / Uzbek-Chechen (rayman v4.0) / OCR + AI Review / LILT Create (Gemma)

AWS

g6.24xlarge

4 L4 GPUs / 96vCPU / 384G RAM / 2TB / 96 GiB VRAM

1

Hour

$6.68

$3,404.58

Batch / Whisper (ASR)

Total Cost

All v4 lang-pairs with standard service

$17.69

$9,168.15

[1] Price calculated for plan: “On-Demand Instances”[2] Price calculated for plan: “Compute Savings Plans” (1 year)[3] Price calculated for plan: “EC2 Instance Savings Plans” (1 year)[4] The standard services require 3 × L4 GPUs (or other 24 GB cards with compute capability 8.0 or higher). T4 GPUs cannot run the translation inference server (rayman), which requires an Ampere-or-newer GPU with at least 22 GB of VRAM; T4 remains valid for the batch worker and ASR. AWS does not currently offer an instance type with exactly 3 x L4 GPUs, so g6.12xlarge (4 × L4) is recommended as the closest available option. Please choose the available instance type best suited to your infrastructure.[5] RDS is priced at On-Demand; the EC2 Instance Savings Plan does not apply to Amazon RDS. A separate RDS Reserved Instance (1 year) would lower this line item.Sources:Recommended EKS deployment
This EKS table follows the Standard tier of the Recommended Production Configurations table above: the platform runs on two GPU nodes (g6e.12xlarge for Translate, Uzbek-Chechen rayman v4.0 and OCR; g6.24xlarge for Batch and Whisper/ASR), backed by an RDS database node — no separate CPU worker node is required at this tier.

  • AI Review & LILT Create — Served by the same Gemma deployment as OCR on the g6e.12xlarge node; no additional instance is required.
  • OCR (Gemma) and ASR (Whisper) — In this configuration these are co-located on the standard GPU nodes above and do not require additional instances. ASR with the Sorani (Kurdish) model needs 25 GB of VRAM on one node and has no CPU fallback.
  • Uzbek and Chechen Language Support — Provided through the neural v4.0 pipeline in rayman (which replaces the former dedicated Emma model), co-located on the g6e.12xlarge node.

GCP

Last Revised: 23 Jul 2026

GCP

Platform

Resource Identifier

Specifications

Quantity

Unit

Unit cost [1]

Cost [1] per month

Cost [2] per month

Notes

GCP

n1-standard-4

4vCPU / 15G RAM / 500GB

1

Hour

$0.19

$138.70

$87.38

K8s Master

GCP

n2-highmem-64

64vCPU / 512G RAM / 1TB

2

Hour

$8.38

$6,120.90

$3,856.01

K8s Nodes

GCP

g2-standard-48

4 L4 GPUs [4] / 48vCPU / 192G RAM / 2TB / 96 GiB VRAM

1

Hour

$4.00

$2,921.24

$2,445.43

Standard Services (Translate, Batch)

Total Cost

All v4 lang-pairs with standard service

$12.57

$9,180.84

$6,388.82

GCP

g2-standard-24 (Optional)

2 L4 GPU / 24vCPU / 96G RAM / 2TB / 48 GiB VRAM

1

Hour

$2.00

$1,460.58

$1,222.75

OCR + AI Review / LILT Create (Gemma)

Total Cost (incl. optional services)

Standard service + OCR/AI Review/Create

$14.57

$10,641.42

$7,611.57

[1] Price calculated for plan: “On-Demand Instances”[2] Price calculated for plan: “Committed use discount options” (1 year)Source: GCP Pricing Calculator EstimateOptional GPU-Accelerated Services
In addition to the standard deployment, several optional services require extra GPU resources:
  • OCR, AI Review & LILT Create — All three are served by one Gemma (vLLM) deployment, which needs 45 GB of VRAM on a single node (one 48 GB card, or two 24 GB cards in the same server) and 80 GB of node RAM, as shown in the optional row above. Enabling AI Review or LILT Create on top of OCR adds no GPU.
  • Uzbek and Chechen Language Support — Provided through the neural v4.0 pipeline in rayman (which replaces the former dedicated Emma model). When enabled, it requires one additional L4 or A10 GPU with at least 24GB VRAM; when disabled, rayman runs on a single GPU.
  • ASR (Automatic Speech Recognition) — GPU-accelerated ASR requires one additional T4 GPU with 16 GB VRAM. By default, it runs on CPU. ASR with the Sorani (Kurdish) model needs 25 GB of VRAM on one node (two 16 GB or two 24 GB cards in the same server, or one 48 GB card) and has no CPU fallback.

Microsoft Azure

Last Revised: 23 Jul 2026

Microsoft Azure

Platform

Resource Identifier

Specifications

Quantity

Unit

Unit cost [1]

Cost [1] per month

Cost [2] per month

Notes

Azure

D3 v2

4vCPU / 14G RAM / 512GiB

1

Hour

$0.29

$213.89

$147.58

K8s Master

Azure

E64as v5

64vCPU / 512G RAM / 512GiB

2

Hour

$7.23

$5,279.36

$3,579.93

K8s Nodes

Azure

Standard_NV72ads_A10_v5

2 A10 GPU / 72vCPU / 880G RAM / 2880GiB/ 48GB VRAM

2

Hour

$13.04

$9,519.20

$7,923.78

Standard Services (Translate, Batch)

Total Cost

All v4 lang-pairs with standard service

$20.56

$15,012.45

$11,651.29

Azure

Standard_NV72ads_A10_v5 (Optional)

2 A10 GPU / 72vCPU / 880G RAM / 2880GiB/ 48GB VRAM

1

Hour

$6.52

$4,759.60

$3,961.89

OCR + AI Review / LILT Create (Gemma)

Total Cost (incl. optional services)

Standard service + OCR/AI Review/Create

$27.08

$19,772.05

$15,613.18

[1] Price calculated for plan: “Pay as you go”[2] Price calculated for plan: “1 year savings plan”Source: https://azure.com/e/41227c0f7aed4a208bbe74cacb734accOptional GPU-Accelerated Services
In addition to the standard deployment, several optional services require extra GPU resources:
  • OCR, AI Review & LILT Create — All three are served by one Gemma (vLLM) deployment, which needs 45 GB of VRAM on a single node (one 48 GB card, or two 24 GB cards in the same server) and 80 GB of node RAM, as shown in the optional row above. Enabling AI Review or LILT Create on top of OCR adds no GPU. The A10 is an Ampere card (compute capability 8.6): it runs Gemma at reduced throughput and is not enabled by the default installer, so confirm with LILT before sizing on the Standard_NV72ads_A10_v5 row above, or choose an Ada-class SKU.
  • Uzbek and Chechen Language Support — Provided through the neural v4.0 pipeline in rayman (which replaces the former dedicated Emma model). When enabled, it requires one additional L4 or A10 GPU with at least 24GB VRAM; when disabled, rayman runs on a single GPU.
  • ASR (Automatic Speech Recognition) — GPU-accelerated ASR requires one additional T4 GPU with 16 GB VRAM. By default, it runs on CPU. ASR with the Sorani (Kurdish) model needs 25 GB of VRAM on one node (two 16 GB or two 24 GB cards in the same server, or one 48 GB card) and has no CPU fallback.