Inference Capital Markets
Lucas Tcheyan (Galaxy Research) · July 13, 2026
The Thesis: Inference Is the New Oil
For the past three years, the AI narrative has been dominated by training runs. GPT-4, Claude 3, Gemini Ultra — every headline was about how many FLOPs went into the latest frontier model. Investors piled into GPU clusters, hyperscalers broke ground on data centers the size of small cities, and the assumption was that whoever trained the biggest model won.
That assumption is flipping.
Inference — the act of running a trained model to generate an output — has overtaken training as the dominant consumer of GPU cycles. By Galaxy's estimates, inference now accounts for roughly 60% of all AI compute demand, and that number is projected to hit 80% within eighteen months. The implications are profound, not just for hardware markets but for the financial infrastructure needed to support them.
Why Inference Changes the Economics
Training is a capital expense. You raise a large round, buy or rent a cluster, run it for weeks or months, and produce a model. The cost is concentrated, the timeline is predictable, and the output is a single asset.
Inference is an operating expense. Every API call, every chat completion, every image generation consumes compute in real time. The cost is distributed across millions of users and thousands of applications. It behaves like a utility — you pay for what you use — but the underlying infrastructure requires the same upfront capital as training.
This mismatch creates a financing problem. If you are a GPU operator, you need to pay for hardware before you see a single dollar of inference revenue. If you are a developer, you want access to compute without committing to a long-term lease. Traditional finance has tools for this — futures contracts, asset-backed lending, capacity auctions — but they are slow, opaque, and designed for markets that operate on quarterly cycles, not millisecond ones.
Onchain Markets for Compute
This is where crypto-native financial primitives enter the picture. The paper argues that the same infrastructure that powers decentralized finance — programmable tokens, automated market makers, onchain credit — can be repurposed to create liquid markets for inference compute.
The core idea is straightforward:
- Tokenized GPU capacity: A hardware operator mints a token representing a unit of future compute (e.g., one hour of H100 inference time in Q3 2026). The token is backed by a specific machine or cluster, with onchain attestations proving the hardware exists.
- Forward markets: Buyers purchase these tokens at a discount to spot price, locking in compute costs and hedging against demand spikes. Sellers get upfront capital to finance hardware deployment.
- Secondary trading: Tokens trade on decentralized exchanges, creating a price signal for compute that is continuous, transparent, and global. The market discovers the time value of GPU cycles.
This is not theoretical. Several projects are already experimenting with this model. The paper cites Akash Network's compute marketplace, Spheron's GPU tokenization, and newer entrants like Exabits and Clover, which are building the settlement layers for these markets.
The Commodification of Intelligence
The paper's most provocative claim is about what happens next. If inference becomes a traded commodity with transparent pricing and liquid forward markets, the marginal cost of AI intelligence drops toward the cost of compute. The model itself — the weights, the architecture, the training data — becomes less important than the ability to access and deploy it efficiently.
This is already visible in practice. Open-weight models like Llama 3 and Mistral are closing the gap with proprietary frontier models. The differentiation is shifting from what the model knows to how cheaply and reliably you can run it. Inference capital markets accelerate this shift by making compute a predictable, financeable input rather than a scarce, volatile one.
Risks and Open Questions
The paper does not shy away from the challenges:
- Hardware attestation: How do you verify that a token is actually backed by a working GPU? Without robust onchain verification, the market is vulnerable to fraud.
- Demand forecasting: Inference demand is spiky and application-dependent. A chat bot uses compute differently than a video generation model. Forward contracts need to account for this heterogeneity.
- Regulatory uncertainty: Compute is not a regulated financial asset today, but as these markets grow, they will attract scrutiny. The line between a compute pre-sale and a security is not clear.
- Latency requirements: Onchain settlement introduces latency that may be unacceptable for real-time inference workloads. The paper suggests layer-2 solutions and optimistic settlement mechanisms as mitigations.
Why This Matters Now
The timing is driven by two converging trends. First, GPU supply is finally loosening after two years of shortages, which means the marginal value of compute is shifting from access to price. Second, the crypto infrastructure — L2 scaling, zk-proofs, oracles — has matured to the point where complex financial instruments can be deployed onchain without the friction that doomed earlier attempts.
If inference capital markets mature, they do not just change how we pay for AI. They change who can build it. A developer in Lagos or Hanoi gets the same access to compute as a well-funded startup in San Francisco, priced by a global market rather than a bilateral negotiation with a cloud provider. That is a world worth building toward.