Princeps is the independent risk and grading layer for compute — call it a credit agency, or even the gold standard, for compute. We price the risk nascent compute markets have so far overlooked, modelling whether a given operator will actually deliver the capacity and uptime it sells. Fundamentally, our thesis is that any GPU-backed loan is only as good as the hardware securing it and the operator running it.
Since last year, compute markets have boomed. Compute now has a price, venues, and a path to physical settlement. What it does not yet have is a deliverable standard (a measure of whose GPU-hours are good delivery and whose are not). Today a top-tier operator and an unreliable one look identical on paper and are procured against the same undifferentiated benchmark. This is more a trust problem than a pricing issue. Historically, every commodity that became tradable first grew an independent layer certifying that what was delivered matched what was promised (gold has the LBMA Good Delivery List, grain has licensed graders and certified warehouses, oil has independent inspection at the delivery point). Liquidity will always follow trust.
We're confident the market is moving in this direction already. Compute is developing observable reference prices through indexes and now futures listings. But a financial contract, as one recent analysis of the compute-offtake market puts it, "cannot make a delayed data centre open on time, secure additional power, repair a network failure, or turn one generation of GPU into another."17 While a price curve prices the market, it still leaves untouched the basis between a reference price and the compute a specific operator actually delivers (the configuration, delivery, service-level and counterparty risk) that is still bundled and unpriced inside private capacity agreements today. Whoever takes on delivery of that compute, whether the platform reselling it, the lab training on it, or the lender financing it, inherits that risk, not the operator. Princeps prices it. As financing and procurement grow more modular, lenders already separate market-price, utilisation, basis, operating and counterparty risk. Independent of any one index or marketplace, Princeps provides the independent grade for compute based on reliability, capacity and delivery.
Founding Team
Backed by Y Combinator and angels and advisors from Standard Intelligence, SF Compute, Kojo, Palantir and 8VC.
USD.ai has built fast, GPU-backed credit that clears in days against hardware tokenised into on-chain warehouse receipts under CALIBER. What is yet to be built is an independent, continuous verification layer: proof that the collateral is physically present, serialised and live, and an independent grade of whether the operator behind it will keep generating the uptime that services the loan. Princeps proposes exactly that layer, in partnership with USD.ai.
Why does this matter for USD.ai? USD.ai originates credit against GPUs held under a datacenter bailment and tokenised into warehouse receipts (GWRTs). Risk is assessed at origination but left to drift post-financing. Princeps solves two questions. First, does the collateral actually exist and stay live on-site — a borrower can move, resell, double-pledge or power down hardware between inspections. Second, will the operator keep delivering — reliability can decay long before a missed payment reveals it. In a default, the remedy is physical repossession of hardware you have never independently verified is there. That risk ultimately lands on sUSDai depositors, who take it today largely on trust.
What we are proposing? One integration that provides both continuous proof-of-collateral and an independent operator grade. Every GWRT gets a live attestation that its serialised hardware is present and online; every borrower gets a Collateral & Credit Grade that maps to an LTV / haircut and an early-warning signal on SLA breach. All of this is auditable, so sUSDai depositors and USD.ai's own investors can verify that every liability is backed by live hardware.
How do we grade? The grade composes two factor groups: the site's physical resilience (from the Princeps database) and the operator's observed collateral & delivery record (from the bailment, loan-servicing history and on-site attestation) into a single score.
| Governed factor | Score |
|---|---|
| On-site generation & storage 11 | 8 / 10 |
| Grid-interconnection age & capacity 13 | 7 / 10 |
| Cooling & thermal headroom | 8 / 10 |
| Chip inventory match | 9 / 10 |
| Permits & build quality | 7 / 10 |
| Governed factor | Score |
|---|---|
| Serial-number & liveness attestation | 6 / 10 |
| Spin-up success rate | 7 / 10 |
| Order-fulfilment history | 6 / 10 |
| SLA-breach & payment-coverage signal | 5 / 10 |
Structural and External Reliability share one evidence-weighting framework — they differ only in the factors, rules and weights. Scores run 1–10 (neutral \(N=5\)); every score also carries an evidence-support value \(S\in[0,1]\) that governs how far it can move from neutral. The number is only as strong as the evidence behind it.
Each reviewed record earns a quality value as a weighted geometric mean across five dimensions, so a record must perform on all of them to score highly:
\(A\) source authority · \(R\) facility relevance · \(D\) measurement directness · \(F\) freshness · \(C\) extraction confidence.
Records from the same document or dataset form one cluster, so duplicates can't inflate support. A cluster's support is its best record; its score is the quality-weighted mean:
Independent clusters corroborate one another, and disagreement between them is penalised:
Support decides how far the observed score \(O_f\) can travel from neutral. Thin evidence stays near neutral; a missing factor stays at 5 rather than being dropped or renormalised away:
Factors roll up into dimensions and dimensions into the overall score, each on fixed weights; the displayed percentage is simply \(100\times\) support:
Evidence support is not the share of documents reviewed — it measures how much of the fixed methodology is genuinely supported once authority, relevance, directness, freshness, extraction confidence, duplicate sources and disagreement are all accounted for.
| Evidence support | Classification |
|---|---|
| Below 30% | Insufficient evidence |
| 30% – 60% | Provisional |
| 60% – 80% | Supported |
| 80% or higher | Strongly supported |
Worked example — London (Docklands): overall reliability 5.49 / 10 at 36.5% evidence support. From the Princeps Neocloud Reliability Scoring Methodology (Hector, Frøyland Moe, Trofimova, 2026).
Princeps already grades neocloud facilities in production, from reviewed engineering and public-market evidence, under a governed, versioned methodology. Below is the live workspace applied to operators across a cross-section of London facilities.
A GWRT records that a borrower pledged a specific fleet of GPUs, held on-site under a bailment. What it does not record, after issuance, is whether those exact serial numbers are still racked, still powered, and still generating the revenue that services the loan. The underwriting file at origination is a snapshot, while collateral progressively deteriorates and performs unreliably. USD.ai originates credit only against hardware that is "active, on-site, and operating under a valid datacenter bailment"6, but the open question of reliability still remains.
| Facility | Collateral | LTV | Health | APR | Serials ✓ | Liveness | Princeps |
|---|---|---|---|---|---|---|---|
| Operator A | 512× H100 | 70% | Healthy | 14.2% | 508 / 512 | 99.2% | Verified |
| Operator B | 256× H100 | 55% | Watch | 16.8% | 198 / 256 | 81.0% | Elevated |
Ownership on paper is not hardware on-site. A GWRT records who owns the GPUs, not whether they are still racked and running — so on-chain, a depleted fleet looks identical to a full one.
What separates a good operator from a bad one is not whether hardware fails (which it inevitably will) but how well that run is contained. In one study, Meta held >90% effective training time because only 3 of 419 incidents needed manual intervention.7 A weaker operator can run up to 10.7% lost compute. The failure waterfall scales with cluster size, meaning that the larger the reserve, the more a run depends on the operator's containment and rerouting capabilities.
Princeps reconstructs the loss history for neoclouds and compute providers from zero to one. Our fundamental belief is that risk must be assessed at origination, across both the physical and performance layers, to be meaningfully quantified. This is the single most mispriced risk in GPU-backed lending right now. Uptime Institute tiers measure redundancy, not AI-readiness (built for 5–10 kW racks when a GB200 rack exceeds 120 kW) and certification is static.10 SLAs cap the remedy at a service credit, usually a 10% credit on a ~$3 instance is about 30 cents, against outages averaging near $1 million.11 The most serious independent effort is ClusterMAX. SemiAnalysis is both the strongest proof that reliability varies enormously and that an independent grade has real traction.
SemiAnalysis's ClusterMAX has been the de-facto GPU-cloud reliability benchmark but only CoreWeave reached their Platinum ranking.12 SemiAnalysis's own verdict: "the bar across the GPU cloud industry is currently very low." The gap is not the silicon. On identical hardware, ClusterMAX measured the difference between a poorly-tuned and a well-tuned operator as roughly 60% vs 98% usable network efficiency. It is the operator, not the GPU, decides whether the compute is usable.
ClusterMAX is a periodic, editorial tier list refreshed quarterly. Princeps extends the same evidence base into a continuous, per-loan risk quantification a lender can act on in real time as a live benchmark. We are independent of any one index or marketplace, allowing us to collect data across every operator, site and venue. Princeps fills this intelligence for both the operator and the lender financing it.
Reliability can be inferred from the physical and structural signals around a provider, and confirmed by whether the provider actually delivered against real orders. Princeps grades ~940 operators into five bands. Crucially, the baseline grade requires no probes inside an operator's stack. It is built from external data plus the bailment and loan-servicing record, and is deepened, at the operator's option, by device-level attestation of the pledged serials.
| Powered by | Princeps Grade | Tier | Delivery confidence |
|---|---|---|---|
| Nebius | Verified | Gold | |
| Crusoe | Verified | Gold | |
| Lambda | Reliable | Silver | |
| Vultr | Reliable | Tier III | |
| DataCrunch (Verda) | Watch | Bronze | |
| Massed Compute | Elevated | Underperformer |
Every grade resolves to a dollar figure a credit team can underwrite against. Dial in the facility (fleet, size, term, checkpoint interval) and Princeps runs the Monte-Carlo model to translate the operator's reliability into expected and tail loss on that specific loan.
Sample demonstration:
| Grade | Reliability 30-day delivery |
Serviced value delivered |
Loss · expected P50 |
Loss · tail P99 |
Failure exposure | Collateral confidence |
|---|
The loss figures are derived from a Monte-Carlo simulation of the facility you set above. Each of 4,000 trials draws on both hardware failures9 and a severity per failure (how much revenue-generating uptime is lost and how long recovery takes, scaled by the operator's containment — the grade). The result is a full distribution of serviced value. The gap between the expected outcome and the unlucky tail (P99) is the loss a lender carries silently today, and exactly what a collateral & credit grade lets USD.ai surface, quantify and price into LTV.
4,000 simulated facilities · Verified operator vs Elevated operator — same face-value collateral
P50 (expected) and P99 (tail) revenue-service loss, $ — grade sourced from each operator's tier
Model: failures ~ Poisson(N·24·T / 50,000); per-failure lost cluster-time ~ checkpoint-interval rework + recovery, exponential severity, containment scaled so each grade's mean matches its measured goodput-loss band [SemiAnalysis/Nebius]. Illustrative — for structure, not a quote. Synchronous training assumed (a failure stalls the whole cluster to the last checkpoint).
We joined six operators representative of a GPU-backed loan pool to the Princeps database and pulled each one's independent tier, ownership model, chip generation and site footprint.
| Provider | Princeps grade | ClusterMAX tier (according to SemiAnalysis) | Usable goodput | Wasted · expected (P50) | Wasted · tail (P99) | Site evidence |
|---|
| Operator | Collateral | Location | Princeps Grade | LTV | |
|---|---|---|---|---|---|
| Operator B | ATTESTED | US | Verified | 75% | View grade |
| Operator D | ATTESTED | Finland | Reliable | 65% | View grade |
A grade fundamentally sets out how much can be safely advanced against a given fleet, because a stronger operator's collateral is worth more against the same hardware. It catches collateral that is quietly deteriorating long before a payment is missed and sizes how much loss the book is actually carrying, so depositors can see the buffer that protects them. Princeps also separates the true cost of capital from the uncertainty premium a lender pays simply for not knowing.
Grounded in USD.ai's own figures, on an identical H100 fleet, the operator grade moves the defensible advance rate from the low-40s to USD.ai's full 70% tier — a swing of roughly 25 points in how much capital the same collateral safely supports, purely on who is running it. It moves modelled annual loss by close to an order of magnitude, from a few tenths of a percent under a Verified operator to several percent under an Elevated one. Carried across the $205M deployed book, grading that spread through the haircut roughly halves expected loss and materially thins the tail. On the pricing side, most of a ~15% borrower rate over the risk-free base is not compensation for credit risk at all, it is an uncertainty premium paid for underwriting blind. An independent grade converts that premium directly into cheaper capital for reliable operators and wider, more durable margin for the protocol.
Advance rate (bars) and modelled annualised expected loss (line) by grade. LTV anchored to USD.ai's liquid-GPU tier (70% cap); loss scaled to SemiAnalysis goodput-loss bands.2
Receipt face value vs continuously-attested live GPUs and liveness over 120 days; failure cadence per Epoch AI / Meta Llama 3.7 The day-90 inspection lags the day-60 covenant breach by a month.
Expected (P50) and tail (P99) loss on USD.ai's $205M deployed book1 — flat LTV vs grade-driven haircuts.
A typical GPU-loan APR (USD.ai Proof of Reserves, 14.2–16.8%)6 split into base rate and the uncertainty premium a grade compresses.
Illustrative models for structure, not quotes. LTV and expected loss are functions of the operator grade; portfolio figures apply grade-driven haircuts across a representative pool; the cost-of-capital view decomposes a typical GPU-loan APR into a base rate and the uncertainty premium an independent grade compresses — the buffer, and the margin, that stand between the loan book and sUSDai depositors.
That a grade transforms a credit market is one of the better-established results in economics. Akerlof's 1970 "market for lemons" showed that when buyers can't observe quality they pay only for average quality, good sellers exit, and the market can unravel. The fix is the creation of "counteracting institutions" such as certification and third-party grading.4
| Market | Credibility layer | Measured effect |
|---|---|---|
| Grain, gold | USDA grades / LBMA Good Delivery | Made the asset exchange-tradeable5 |
| Corporate debt | Credit ratings (NRSRO) | Unlocked mandated institutional capital14 |
| Enterprise SaaS | SOC 2 attestation | De facto gate to enterprise procurement15 |
The mechanism is identical: an independent layer converts an unobservable quality attribute into a verifiable, priceable signal — raising prices for good sellers, raising transaction probability, widening participation, and shrinking the pooling that drives adverse selection.
The cost of capital falls for good operators. A grade is most valuable exactly where trust is scarcest and borrowers are otherwise indistinguishable, as in GPU-backed lending. Capital shifts to quality, and the depositor pool widens: allocators commit to sUSDai on verifiable, auditable backing, not trust. An independent, quantified collateral & credit grade brings in capital that is currently sitting out for want of exactly this assurance.
Underneath the grade sits a continuous intelligence layer that scores every operator across six dimensions:
Princeps would provide ongoing risk intelligence, collateral verification and credit scoring for USD.ai — a live index across every operator in the loan book. It can run as a standalone underwriting-and-audit service for USD.ai and its curators, or integrate directly into the protocol: grade-driven LTV, the attestation feed, and an investor-facing audit view all drawing on the same index.
The ecosystem already has a partner for every layer except this one. We would slot in as the independent Risk & Collateral-Verification category — the missing counterpart to the data and security partners USD.ai already relies on.
| Ecosystem layer | Today | What it secures |
|---|---|---|
| Market data / oracle | Chainlink · Chronicle | On-chain prices & feeds |
| Security | Cantina · Spearbit · Immunefi | Smart-contract integrity |
| Custody | Fireblocks · Anchorage · Copper | Asset safekeeping |
| Risk & collateral verification | Princeps — proposed | Proof of performance, operator reliability & hardware |
The integration above is the starting point, and we are keen to explore a larger, longer-term collaboration alongside it. USD.ai generates data most of the market cannot see — a continuous record of which operators actually serviced their loans and whose collateral stayed live, drawn from every facility it originates. Combined with the Princeps methodology, that record is the foundation for a jointly-operated Compute Credit Rating. Over time it would position USD.ai as the reference point for GPU-backed credit, much as the Chicago Board of Trade's grading standards became the basis on which the grain market settled.5
The path from independent grade to rating standard runs in three phases:
Princeps operates the rating independently, with USD.ai as the data spine and reference venue. A rating the lender controls is not an independent rating. Independence is what makes the standard credible to depositors, investors and counterparties and it is what underpins protocol solvency and depositor confidence over the long term. The rating improves the more volume USD.ai originates.
1. USD.ai "Lighthouse" 2026 YTD report & live protocol data — TVL, deployed, sUSDai yield, signed term sheets — via Dune (Entropy Advisors) and DefiLlama, 2026.
2. SemiAnalysis (commissioned by Nebius), goodput-loss benchmarking, 2026.
3. USD.ai, protocol risk disclosures on default, repossession & remarketing of collateral, 2025–26.
4. Akerlof, "The Market for 'Lemons'," QJE 84(3), 1970.
5. USDA AMS; LBMA Good Delivery; CME Group (CBOT futures, 1865).
6. USD.ai docs (docs.usd.ai) & "How USD.AI GPU Loans Work": Proof of Reserves, CALIBER framework & GWRTs, non-recourse facilities, asset-level LTV, credit originated only against active, on-site hardware under a valid bailment, 2025–26.
7. Meta, "The Llama 3 Herd of Models," arXiv:2407.21783, 2024.
8. Meta, "Revisiting Reliability in Large-Scale ML Research Clusters," arXiv:2410.21680, 2024.
9. Epoch AI, GPU-cluster failure-scaling model (~1/50,000 GPU-hrs), 2024.
10. American Compute, "Data Center Tiers," 2026; Uptime Institute.
11. Uptime Institute, "Annual Outage Analysis 2024"; "Cloud SLAs punish, not compensate," 2022.
12. SemiAnalysis, "ClusterMAX 2.0," 2025.
13. LBNL, "Queued Up: 2025 Edition," 2025.
14. White, "The Credit Rating Agencies," JEP 24(2), 2010.
15. ControlCase (industry reporting), 2024–25.
16. McKinsey, "The evolution of neoclouds," 2025–26.
17. Dave Friedman, "The Compute Market has Multiple Views on Future Compute Prices," 2026.
18. USD.ai Ecosystem (usd.ai/ecosystem); protocol built by MetaStreet Labs; API at api.usd.ai.
For discussion only; not an offer of insurance or a quotation. Marketplace prices are pulled live and change in real time. The failure simulation is calibrated to measured goodput-loss and failure-rate data and is illustrative of structure, not a quote. Grades are a framework, not a published assessment of any provider.