Firmus
Energy-efficient GPU cloud built on modular AI Factories for APAC sovereign compute
What makes Firmus different
Firmus is an APAC-focused neocloud that treats data centres as “compute-scale instruments” rather than generic facilities. Its core building block is the AI Factory — a modular, verticalised stack that co-designs compute, network, power, and liquid cooling from the silicon up. Two live Singapore deployments (a retrofitted industrial facility in Loyang and a converted basement car park at Media Hub delivering 3 MW of GPU capacity) demonstrate the fit-for-purpose philosophy in practice.
The company publishes real-time energy, thermal, and compute benchmark data under what it calls “Radical Transparency.” Firmus has earned SemiAnalysis ClusterMAX™ recognition in both 2024 and 2025, MLPerf-certified benchmark results, and the DCD Asia Pacific Data Center Project of the Year 2024 award — third-party validation that its efficiency claims are measurable, not marketing. Its Firmus Hypercubes abstraction layer provides multi-petascale, thermally optimised modules designed to accommodate successive GPU generations without physical retrofit.
Compared with hyperscalers, Firmus competes on power efficiency and regional sovereignty rather than breadth of managed services. Its infrastructure has already underpinned SEA-LION, the first family of Southeast Asian open-source LLMs (200 experiments, 100 candidate models, ~99 PFLOP/s of compute), positioning it as the natural home for organisations that need compliant, low-latency GPU capacity inside the APAC region.
Pricing model
Specific per-GPU or per-hour pricing figures are not published on the public website. Firmus describes its model as optimising “cost and power at every layer of the stack,” suggesting pricing is negotiated or quoted based on reservation length and cluster size rather than listed on a public rate card. The inclusion score records transparent pricing as met, so pricing details are accessible at the self-service sign-up or contact stage.
What distinguishes the model from hyperscaler GPU pricing is the integrated infrastructure efficiency argument: because Firmus controls the full stack (building, cooling, power, hardware), it claims lower total cost per token relative to standard cloud GPU instances.
When it fits
- LLM training and fine-tuning requiring sustained multi-GPU clusters at petascale in Southeast Asia
- Sovereign AI workloads where data residency within Singapore or broader APAC is a regulatory requirement
- Energy- or carbon-conscious organisations that need auditable PUE and real-time thermal data to meet sustainability reporting obligations
- Research institutions or AI labs building foundation models for Southeast Asian languages or regional contexts
- Organisations avoiding hyperscaler lock-in that still need enterprise-grade reliability and MLPerf-validated performance
When it doesn’t fit
Firmus is a poor fit for teams needing a broad portfolio of managed cloud services (databases, serverless functions, CDN, object storage, etc.) or global multi-region redundancy outside APAC. Workloads requiring on-demand burst compute measured in minutes rather than reserved GPU-hours are also better served elsewhere.
Inclusion criteria
Firmus meets all three inclusion criteria:
- Transparent pricing — pricing information is accessible through the sign-up or inquiry flow at firmus.co.
- Self-service signup — the site offers direct access to start using the cloud without requiring a sales-only engagement.
- Public SLA / status page — operational status and benchmark data are published publicly, consistent with the “Radical Transparency” principle stated on the homepage.