Learn AI compute, then follow the market

Learn AI compute in plain English.

Free lessons, calculators, and explainers for understanding GPUs, model costs, prompt and context design, cloud capacity, data centers, power constraints, and emerging AI compute infrastructure. Explore 81 published lessons, follow a learning path, or use a calculator to estimate real compute costs.

A different lesson is highlighted in the Morning Brief each day.

The curriculum

9 tracks, plain English

Pick a track and work through it in order, or jump to the topic you need. Track your progress in this browser — no account required. Every lesson is free.

Track 1

AI Compute 101

Start here. GPUs, GPU-hours, model costs, and why compute became a market.

0 of 7 complete

View track →

Track 2

GPU Pricing & Capacity

What H100, H200, and B200 hours cost, and how on-demand, reserved, and spot pricing differ.

0 of 8 complete

View track →

Track 3

Model Costs

What it costs to train a model and to serve one, and why utilization drives the bill.

0 of 7 complete

View track →

Track 4

Prompt & Context Engineering

How prompts, context, retrieval, and agent state shape model quality, token usage, latency, and inference cost.

0 of 2 complete

View track →

Track 5

Model Benchmarks & AI Compute Economics

How benchmark scores connect to token pricing, latency, throughput, and real inference spend.

0 of 23 complete

View track →

Track 6

Power & Data Centers

Power, cooling, networking, memory, and the physical sites where AI compute actually runs.

0 of 17 complete

View track →

Track 7

Compute Market Structure

Neoclouds, capacity markets, and reservations: how compute supply gets priced.

0 of 8 complete

View track →

Track 8

Compute Futures

Forward pricing for compute: futures, forward curves, and forward contracts.

0 of 5 complete

View track →

Track 9

Buyers & Operators

How labs and teams buy, compare, and budget GPU capacity in practice.

0 of 8 complete

View track →

Compute College tracks

Explore the 9 tracks

Once you know the basics, work through pricing, model costs, prompt and context design, infrastructure and power, market structure, compute futures, and the buyer-and-operator playbook. Emerging topics and the calculators sit alongside them.

AI Compute 101

What is AI compute?

AI compute is the hardware, power, networking, and cloud access used to train and run AI models.

Why compute matters

AI capability becomes a market constraint when usable compute is scarce, costly, or slow to deliver.

GPU hours meaning: what is a GPU-hour?

The core unit for calculating AI GPU rental cost, utilization, and workload budgets.

H100 vs H200 vs B200: price, memory, and performance

How NVIDIA accelerator generations compare on workload fit, current GPU-hour pricing, memory, and availability.

What is Model Training Cost?

Model training cost is the compute expense required to teach or improve an AI model from data.

What is Frontier Model Serving Cost?

Frontier model serving cost is the estimated expense of running a leading AI model for users after training.

Why power matters for AI compute capacity

Why electricity, interconnection, and site readiness now constrain GPU deployment.

GPU Pricing & Capacity

NVIDIA H100 Price Per Hour Explained

H100 price per hour is the hourly cost of accessing one NVIDIA H100 GPU for an AI workload.

H200 Price Per Hour Explained

H200 price per hour is the hourly cost of accessing one NVIDIA H200 GPU for AI workloads.

B200 Price Per Hour Explained

B200 price per hour is the hourly cost of accessing NVIDIA B200-generation AI compute capacity.

GPU vs TPU vs Custom ASIC

GPUs, TPUs, and custom ASICs are different kinds of AI accelerator that trade flexibility for efficiency on targeted workloads.

On-Demand vs Reserved vs Spot GPU Pricing

On-demand, reserved, and spot GPU pricing are three ways buyers obtain and pay for AI compute capacity.

What are GPU rentals? AI cloud pricing and capacity signals

How rented accelerator capacity turns GPU access into observable market pricing.

What are spot prices? AI GPU spot market explained

How interruptible GPU capacity prices expose marginal AI compute supply.

What is GPU Cloud Capacity?

GPU cloud capacity is buyer-accessible accelerator supply available for AI workloads through cloud providers.

Model Costs

What is Model Training Cost?

Model training cost is the compute expense required to teach or improve an AI model from data.

What is Frontier Model Serving Cost?

Frontier model serving cost is the estimated expense of running a leading AI model for users after training.

What is Cost per Million Tokens?

Cost per million tokens is how hosted AI APIs price inference — usually with input and output tokens priced separately.

What is GPU Utilization?

GPU utilization measures how much paid accelerator capacity is actively doing useful work.

What is Model FLOPs Utilization (MFU)?

Model FLOPs Utilization (MFU) measures how much of a GPU's theoretical compute a job actually uses; goodput counts useful work delivered.

How Fast Does an H100 Depreciate?

GPU depreciation spreads an accelerator's purchase cost over its useful life, which drives the real cost of every GPU-hour.

How to Estimate Monthly AI Compute Burn

Monthly AI compute burn measures recurring spending on the capacity used to train, fine-tune, experiment, and serve models.

Prompt & Context Engineering

What is prompt engineering?

How clear instructions, examples, constraints, and output formats help an AI model complete a task.

What is context engineering?

How AI systems select and manage the information a model receives before making a decision.

Model Benchmarks & AI Compute Economics

What are AI model benchmarks?

Learn what AI model benchmarks measure, where they mislead, and why benchmark results can become compute demand and cost signals.

How are AI model benchmarks calculated?

AI model benchmarks compare models on fixed tasks, but their scores only become useful for AI compute buyers when read with cost, latency, and token use.

Why AI model benchmarks can be misleading

Learn why AI benchmark scores can mislead buyers when they hide prompt setup, retries, tool use, latency, token usage, and model serving cost.

How to compare model quality vs cost

Learn how to compare AI model benchmark performance with token pricing, latency, throughput, and cost per useful result.

Benchmark score vs production cost

Learn why higher AI benchmark scores may not lower production cost, and how token usage, latency, retries, and context size affect serving spend.

How to estimate cost per completed AI task

Learn how to estimate the full cost of an AI task, including input tokens, output tokens, retries, tool calls, latency, and model selection.

Model latency explained

Learn what AI model latency means, why it matters for production workloads, and how it connects to model serving cost and infrastructure capacity.

Tokens per second explained

Learn what tokens per second means, how model throughput affects AI applications, and why throughput matters for AI compute capacity planning.

Context window explained

Learn what an AI model context window is and how longer context affects token cost, memory, latency, and model serving economics.

What is a coding benchmark?

Learn what AI coding benchmarks measure and why coding-agent benchmarks matter for inference demand, model serving cost, and AI compute capacity.

What is SWE-bench?

Learn what SWE-bench measures, why it matters for AI coding agents, and how software-engineering benchmarks connect to AI compute demand.

What is LiveCodeBench?

Learn what LiveCodeBench measures, why fresh coding tasks matter, and how contamination-resistant coding benchmarks affect AI model evaluation.

What is Terminal-Bench?

Learn what Terminal-Bench measures and why terminal-based AI agent benchmarks matter for token usage, latency, and AI compute demand.

Claude Opus 4.8 benchmark explained

Read Claude Opus 4.8 benchmark claims as AI compute economics evidence: capability-per-dollar, effort settings, fast mode, agent workloads, and serving demand.

What is Claude Mythos Preview?

Claude Mythos Preview is an unreleased Anthropic frontier model used in Project Glasswing for defensive cybersecurity work.

What is GPQA Diamond?

Learn what GPQA Diamond measures, why expert science reasoning benchmarks matter, and how they connect to frontier AI compute demand.

What is MMLU-Pro?

Learn what MMLU-Pro measures, how it differs from older academic benchmarks, and why benchmark difficulty matters for AI model evaluation.

What is Humanity’s Last Exam?

Learn what Humanity’s Last Exam measures and why frontier academic benchmarks matter for model capability claims and AI compute demand.

What is a reasoning benchmark?

Learn what AI reasoning benchmarks measure and how reasoning scores connect to model serving cost, latency, and frontier AI compute demand.

What is an agent benchmark?

Learn what AI agent benchmarks measure and why agentic workflows can drive higher token usage, latency, retries, and AI compute demand.

How model releases affect AI compute demand

Learn how new AI model releases can change inference demand, training demand, token usage, cloud GPU capacity, and the AI compute market.

Why output tokens cost more than input tokens

Learn why output tokens usually cost more than input tokens and how generation cost affects model serving economics, AI agents, and inference spend.

Why Reasoning Models Cost More to Serve

Reasoning models generate long chains of thought before answering, multiplying output tokens — and output tokens drive inference cost.

Power & Data Centers

Why power matters for AI compute capacity

Why electricity, interconnection, and site readiness now constrain GPU deployment.

What is an AI data center? Power, cooling, and GPU capacity

The facility stack that turns accelerators into operating AI compute supply.

Why cooling matters

Cooling removes the heat created by dense AI hardware so a facility can deliver usable compute safely.

Why networking matters for AI clusters and training cost

How interconnect quality turns GPU count into useful clustered compute.

Why memory matters for AI accelerators and HBM supply

How high-bandwidth memory affects model fit, GPU value, and accelerator availability.

What is an AI Cluster?

An AI cluster is a connected system that turns many GPUs and supporting infrastructure into usable model-training or serving capacity.

What is NVLink?

NVLink is a high-speed GPU connection technology that helps accelerators coordinate work inside AI systems.

What is InfiniBand?

InfiniBand is high-performance networking used to connect servers in many large AI clusters.

What is NVL72? Scale-Up vs Scale-Out

NVL72-style rack systems link many GPUs into one; it illustrates scale-up (bigger tightly-coupled units) versus scale-out (more networked units).

What is Liquid Cooling?

Liquid cooling removes heat from dense AI hardware so more compute can operate reliably in a facility.

What is Data Center Interconnection?

Data center interconnection links AI capacity to networks, clouds, data sources, and buyers who need to use it.

PUE Meaning and Power Usage Effectiveness Formula

Power Usage Effectiveness measures how much facility electricity is required to deliver useful IT power.

Power Sourcing for AI: PPAs and Behind-the-Meter

AI data centers secure electricity through PPAs, behind-the-meter generation, and firm low-carbon sources like nuclear and SMRs.

What is a Megawatt of AI Compute?

A megawatt of AI compute is a power-based way to describe possible data-center and accelerator capacity.

What is a Data Center Interconnection Queue?

A data center interconnection queue is the waiting line for large facilities seeking electrical grid connection.

What is AI Data Center Cooling Density?

AI data center cooling density describes the heat-removal capacity required by concentrated GPU racks.

What is High-Bandwidth Memory (HBM)?

High-bandwidth memory is fast memory located near advanced accelerators to keep AI workloads supplied with data.

Compute Market Structure

What is a neocloud? Meaning, examples, and GPU capacity

How compute-first cloud operators sell GPU clusters, reservations, and AI capacity.

What is a Compute Capacity Market?

A compute capacity market organizes access to AI compute as priced, reserved, allocated, or future-delivered capacity.

What is GPU Cloud Capacity?

GPU cloud capacity is buyer-accessible accelerator supply available for AI workloads through cloud providers.

What is a Compute Reservation?

A compute reservation secures defined GPU or accelerator capacity for a buyer over an agreed period.

What are GPU rentals? AI cloud pricing and capacity signals

How rented accelerator capacity turns GPU access into observable market pricing.

What are spot prices? AI GPU spot market explained

How interruptible GPU capacity prices expose marginal AI compute supply.

What is GPU-Backed Financing?

GPU-backed financing is borrowing against GPUs or their rental contracts to fund AI infrastructure without paying the full cost upfront.

What is Sovereign AI Compute?

Sovereign AI compute is AI capacity a country controls within its own borders and jurisdiction, reducing dependence on foreign infrastructure.

Compute Futures

What are compute futures? AI GPU forward pricing explained

How forward prices turn future GPU capacity into a planning and market signal.

How to read a compute forward curve for AI GPU capacity

How curve shape signals expected tightness, relief, or uncertainty in future compute supply.

What is a Compute Forward Contract?

A compute forward contract agrees today on a price for defined compute capacity delivered at a future time.

What is a Compute Capacity Market?

A compute capacity market organizes access to AI compute as priced, reserved, allocated, or future-delivered capacity.

What is a Compute Reservation?

A compute reservation secures defined GPU or accelerator capacity for a buyer over an agreed period.

Buyers & Operators

How AI Labs Buy Compute

AI labs secure usable capacity through rentals, reservations, cloud agreements, owned clusters, and strategic infrastructure deals.

How to Compare GPU Cloud Quotes

Comparing GPU cloud quotes means normalizing rate, capacity quality, access terms, and expected completed-workload cost.

GPU Cloud Quote Comparison Checklist

Eleven concrete fields to compare on every GPU cloud quote before signing.

How to Estimate Monthly AI Compute Burn

Monthly AI compute burn measures recurring spending on the capacity used to train, fine-tune, experiment, and serve models.

When to Use Spot GPUs vs Reserved Capacity

Spot GPUs suit flexible work; reserved capacity suits predictable or critical AI workloads that need dependable access.

What is a Neocloud Service Level Agreement?

A neocloud service level agreement defines reliability and remedy terms for specialist GPU-cloud capacity.

What is AI Compute Procurement?

AI compute procurement is the disciplined process of sourcing and managing accelerator capacity for AI workloads.

What is GPU Utilization?

GPU utilization measures how much paid accelerator capacity is actively doing useful work.

Tools & Calculators

GPU-Hour Cost Calculator

Estimate AI compute cost from GPU price, runtime, utilization, and overhead.

AI Training Cost Calculator

Estimate a training-run budget using GPU-hours and operating assumptions.

Model Serving Cost Calculator

Estimate recurring inference cost from usage and capacity needs.

Reserved vs On-Demand Calculator

Compare committing to reserved GPU capacity against paying on-demand for a workload.

API vs Self-Hosted Calculator

Compare paying per token for a hosted API against running the model on your own GPUs.

Emerging Topics

Claude Opus 4.7 benchmark explained

Read Claude Opus 4.7 benchmark claims as AI compute economics evidence: capability, token pricing, workload fit, and likely inference demand.

What are Prometheus and Hyperion? Meta AI campus capacity explained

How multi-gigawatt AI campuses translate power plans into future compute supply.

What is Colossus? xAI supercomputer capacity explained

Why a large GPU cluster is a compute-market signal only when powered, cooled, and usable.

What is Project Rainier? AWS Trainium capacity explained

How custom AI silicon can expand or redirect demand for GPU-equivalent compute.

What is Stargate? AI data center capacity project explained

How a mega-scale AI infrastructure buildout can affect future compute supply.

What is Terafab? AI chip supply and compute capacity explained

Why a proposed fab or capacity project matters only when milestones become deliverable supply.

Put it to work

Turn the lesson into a number

Use the GPU-Hour Cost Calculator, AI Training Cost Calculator, or Model Serving Cost Calculator to estimate real compute costs from your own inputs.

Open the calculators

Keep up with the market

Follow the market after the lesson

Read the ComputeTape Morning Brief for daily AI compute pricing, power, capacity, and infrastructure signals — plus a different Compute College lesson highlighted each day.

Read the Morning Brief

Contact Compute College

Help improve Compute College

Send lesson ideas, corrections, source material, or questions about AI compute education to the editorial team.

editorial@computetape.com →