Sell your self-hosted AI model to agents

If you run a model on your own hardware with Ollama, llama.cpp, or vLLM, you can sell it per query: keep the model server private on localhost, put an L402 gate in front of it, expose only the gate over HTTPS, and list it on Amazap. Any agent with a Lightning balance can then pay a few sats, send a chat request, and get the completion back, with no account and no API key to hand out. Your GPU earns while it would otherwise sit idle, and your customers are agents you'd never have reached.

Part of How to sell your API to AI agents.

The setup in one picture

  1. Model server on localhost (Ollama, llama.cpp's llama-server, or vLLM).
  2. L402 gate in front of it: Lightning Labs' Aperture, the Amazap starter kit, or your own code. It issues the invoice, checks the payment, and forwards paid requests.
  3. Public HTTPS that reaches the gate, not the model server: a reverse proxy or a tunnel.
  4. A listing on Amazap, so agents can find and buy it through the Amazap MCP.

Step 1: Run the model server and keep it private

All three servers speak an OpenAI-compatible chat API, which is what most agents already know how to call.

ServerOpenAI-compatible endpointDefault bindNote
Ollama/v1/chat/completions (Ollama docs)127.0.0.1:11434 (Ollama FAQ)Change with OLLAMA_HOST; leave it on localhost when a gate sits in front
llama.cpp (llama-server)OpenAI-compatible API (README)--host defaults to 127.0.0.1--api-key adds a key check; --n-predict caps output tokens
vLLM/v1/chat/completions and more (vLLM docs)Started with vllm servevLLM warns that --api-key only protects /v1, /v2, and /inference, not /invocations, and recommends a reverse proxy

The rule for all three: the internet should never reach the model server directly. Ollama's FAQ shows how to expose Ollama with ngrok or Cloudflare Tunnel; if you sell access, point the tunnel at your L402 gate instead.

Step 2: Put an L402 gate in front

Option A: Aperture (Lightning Labs)

Aperture is Lightning Labs' L402 reverse proxy for REST and gRPC backends. You define services by host and path patterns, give each a backend address and a price in sats, and Aperture issues the challenge and forwards paid requests. A few things from its README and sample config that matter for model hosting:

A trimmed service block, adapted from the sample config (check the sample for every required field):

services:
  - name: "llm"
    hostregexp: '^llm.example.com
  


    pathregexp: '^/v1/chat/completions
  


    address: "127.0.0.1:11434"
    protocol: http
    price: 20

Option B: the Amazap starter kit with Amboss Payments

The Amazap seller starter kit has L402 seller code in Node and Python. In front of a model server, the handler does three things: create an invoice through Amboss Payments, return the L402 challenge, and on a paid request, forward the body to the local OpenAI-compatible endpoint and return the reply. Amboss's create_receive returns the invoice and its payment hash (Amboss docs); payment setup is in Set up Amboss Payments.

The starter kit includes a model proxy for Ollama, llama.cpp server, or vLLM. A paid POST /v1/chat/completions is forwarded to an OpenAI-compatible endpoint on localhost. The proxy caps max_tokens and forces stream to false. An unpaid request, including an empty POST, gets HTTP 402 before any body check.

Set up Amboss Payments

Disclosure: Amazap is an Amboss referral partner. Amazap earns a referral credit if you sign up through this link.

Step 3: Price per query

Amazap lists one price per endpoint: the amount in your invoice. Inference cost varies with output length, so make the price predictable by capping the output:

Aperture's metered bundles are the more precise answer to variable cost, but the Amazap listing flow reads a single per-call price from your invoice, so fixed price per capped query is the fit for listing on Amazap today. Worked pricing examples: How to price an API for AI agents.

Step 4: Run it like a service

Agents favor listings that answer fast and every time. On Amazap, a listing that fails after payment can be paused, and the buyer gets the call credited back.

Step 5: List it on Amazap

Paste the gate's public URL into /sell. Amazap calls it once, unpaid, and expects the L402 challenge; the price comes from the invoice. Make sure the gate answers an unpaid POST with no body with the 402, not a validation error. Then add an example request body (model name, a short prompt, your token cap) so agents can copy it. Walkthrough: List your L402 endpoint on Amazap.

List your endpoint on Amazap

Amazap's catalog already includes L402 image and video models from PayPerQ, so model inference is a category agents shop for. The buyer side is in Agentic payments explained.

FAQ

Can I make money from a self-hosted LLM?

Yes. Put the model server behind an L402 gate, such as Lightning Labs' Aperture or the Amazap seller starter kit, and charge a fixed price in sats per query. List the gated endpoint on Amazap so agents can find and buy it.

Can I sell access to Ollama?

Yes. Keep Ollama on its default localhost bind, put an L402 gate in front of its OpenAI-compatible /v1/chat/completions endpoint, and expose only the gate over HTTPS.

Does Aperture work with Ollama, llama.cpp, or vLLM?

Aperture is a general L402 reverse proxy for REST and gRPC backends, so it can sit in front of any HTTP model server. It needs an lnd node for invoices. Test your configuration against the sample config in the Aperture repo.

How do I price per query when outputs vary in length?

Cap the output tokens on each endpoint and charge a fixed price for that cap. Offer longer outputs as a separate endpoint with a higher price. Aperture also offers metered bundles, but Amazap listings use one per-call price from your invoice.

Is it safe to expose my model server to the internet?

Expose the L402 gate, not the model server. Keep the server on localhost, add its own API key as a second lock where supported, and rate-limit at the gate. vLLM warns that its --api-key does not protect every endpoint.

Do I need a GPU?

No, but speed matters to agents. Smaller models on CPU can work for short tasks. Whatever you run, state the model and limits in your listing.

Can the response stream?

No. Amazap does not pass Server-Sent Events through. It buffers up to 256 KB and stops at 20 seconds, or 90 seconds for document calls. Force stream to false and return one JSON body.

Sources