Sell your self-hosted AI model to agents
If you run a model on your own hardware with Ollama, llama.cpp, or vLLM, you can sell it per query: keep the model server private on localhost, put an L402 gate in front of it, expose only the gate over HTTPS, and list it on Amazap. Any agent with a Lightning balance can then pay a few sats, send a chat request, and get the completion back, with no account and no API key to hand out. Your GPU earns while it would otherwise sit idle, and your customers are agents you'd never have reached.
Part of How to sell your API to AI agents.
The setup in one picture
- Model server on
localhost(Ollama, llama.cpp'sllama-server, or vLLM). - L402 gate in front of it: Lightning Labs' Aperture, the Amazap starter kit, or your own code. It issues the invoice, checks the payment, and forwards paid requests.
- Public HTTPS that reaches the gate, not the model server: a reverse proxy or a tunnel.
- A listing on Amazap, so agents can find and buy it through the Amazap MCP.
Step 1: Run the model server and keep it private
All three servers speak an OpenAI-compatible chat API, which is what most agents already know how to call.
| Server | OpenAI-compatible endpoint | Default bind | Note |
|---|---|---|---|
| Ollama | /v1/chat/completions (Ollama docs) | 127.0.0.1:11434 (Ollama FAQ) | Change with OLLAMA_HOST; leave it on localhost when a gate sits in front |
llama.cpp (llama-server) | OpenAI-compatible API (README) | --host defaults to 127.0.0.1 | --api-key adds a key check; --n-predict caps output tokens |
| vLLM | /v1/chat/completions and more (vLLM docs) | Started with vllm serve | vLLM warns that --api-key only protects /v1, /v2, and /inference, not /invocations, and recommends a reverse proxy |
The rule for all three: the internet should never reach the model server directly. Ollama's FAQ shows how to expose Ollama with ngrok or Cloudflare Tunnel; if you sell access, point the tunnel at your L402 gate instead.
Step 2: Put an L402 gate in front
Option A: Aperture (Lightning Labs)
Aperture is Lightning Labs' L402 reverse proxy for REST and gRPC backends. You define services by host and path patterns, give each a backend address and a price in sats, and Aperture issues the challenge and forwards paid requests. A few things from its README and sample config that matter for model hosting:
- It needs an lnd node. The
authenticatorsection connects to lnd directly (lndhost, TLS cert, macaroon directory) or through a Lightning Node Connect pairing phrase. - Per-service price.
priceis the L402 price in sats, or you can run a separate price server (dynamicprice). - Upstream headers.
headerslets Aperture add a header to every forwarded request, so the model server can keep its own--api-keyas a second lock. - Rate limits. Per-path token-bucket limits, applied per L402 token.
- Metered pricing for inference. Aperture's README notes that "a fixed price per request is the wrong unit for LLM inference" and offers a mode that sells a prepaid bundle and draws it down by the usage the model reports, through a price server called
meterd.
A trimmed service block, adapted from the sample config (check the sample for every required field):
services:
- name: "llm"
hostregexp: '^llm.example.com