Researched Tuesday, October 6, 2026 (PT). Every claim links to its source. Anything I could not confirm is marked Unverified. I didn't sign up for anything or spend money. Prices are in US dollars per 1 million tokens ("per 1M") unless noted.
Bottom line
- What it is: Tinfoil runs open-weight AI models inside hardware-sealed "secure enclaves." Its own SDK checks those enclaves cryptographically before sending anything, so Tinfoil itself can't read your prompts. It's the strongest privacy guarantee of the options compared here (docs, FAQ).
- Is it as smart as Venice? It matches Venice's private models. GLM-5.3 scores 45 on the Artificial Analysis Intelligence Index, against 46 for Grok 4.7, which Venice labels its "most intelligent" model. It is well behind the Claude, GPT and Gemini models Venice resells in "anonymized" mode (Claude Opus 5.5 scores 58) (AA, Venice models API). Use GLM-5.3 as the "smart" slot.
- Filtering: Tinfoil adds no moderation layer to the API (FAQ), but it also has no "uncensored" models like Venice does. Each model keeps its built-in refusals.
- Cost: pay-as-you-go per token, no subscription needed for the API (pricing). A heavy personal user (1,200 messages a month) would pay about $7–$14 a month with the mix recommended below.
- Plug-in difficulty: low. Run
pip install tinfoiland swapOpenAI(...)forTinfoilAI(...). The calls are otherwise the same (Python SDK). - Verdict: Yes, add it as Chance AI's "Private mode" toggle. It's the strongest privacy option, and it's independent of OpenRouter. Keep Venice (or OpenRouter's Venice endpoints) for low-filter chats.
Top 5 Tinfoil models for Chance AI (ranked)
Ranked on four things: intelligence compared with Venice, filtering, speed, and price. Prices come from Tinfoil's official model catalog feed (api.tinfoil.sh/api/config/models, which powers tinfoil.sh/models).
| # | Model (API id) | Best for | Price in / out per 1M |
|---|---|---|---|
| 1 | GLM-5.3 (glm-5-3) |
"Smart" slot. The smartest Tinfoil model (AA index 45), close to Venice's best private model, with a 1M-token context. | $1.80 / $5.75 (cached input $0.45) |
| 2 | DeepSeek V4.1 Flash (deepseek-v4-1-flash) |
Everyday default. Nearly as smart as Gemini 3.8 Flash (AA 39 vs 41), very fast, reads images, 1M context. Marked "experimental" and has had the most outages lately. | $0.65 / $1.45 (cached $0.13) |
| 3 | Kimi K3 (kimi-k3) |
"Max" slot for the hardest questions and image analysis. Highest LMArena score of any Tinfoil model, but slow and expensive. | $4.00 / $20.00 (cached $0.80) |
| 4 | Gemma 4 31B (gemma4-31b) |
Cheap image reading and light chat. The second-least filtered Tinfoil model in Tinfoil's own safety tests. | $0.40 / $1.00 |
| 5 | Llama 3.3 70B (llama3-3-70b) |
"Looser" slot. The least filtered Tinfoil model by a wide margin, but much less smart (AA 8) and pricey for its quality. | $1.75 / $2.75 |
Not in the top 5: gpt-oss-120b is the cheapest at $0.15 / $0.60, but it's the most restrictive model Tinfoil hosts (it complied with only 0.5% of test prompts) and much less smart (AA 12). GLM-5.3 Flash is deprecated and will be removed October 9, 2026 (changelog).
Contents
- What Tinfoil is and how it works
- Models
- Intelligence vs Venice
- Pricing and plans
- API and a Flask code sketch
- Rate limits, reliability, uptime
- Comparison: Tinfoil vs Venice vs Privatemode vs Scaleway
- Downsides and limitations
- Verdict and monthly cost
- Sources
1. What Tinfoil is and how it works
In plain English
Normally an AI provider decrypts your message on its servers, and its staff, its cloud host, or a court order could in principle get at it. Tinfoil runs each model inside a confidential virtual machine. The chip encrypts that machine's memory so the server operator (Tinfoil, or its cloud provider) can't look inside. Before your app sends anything, Tinfoil's SDK asks the chip for a signed "fingerprint" of exactly what software and model weights are running. It compares that fingerprint with the one Tinfoil published publicly from its open-source code. Only if they match does it encrypt your message to a key that exists only inside that enclave (verification docs, enclave primer).
Hardware
- What runs today: "Tinfoil currently runs on AMD SEV-SNP CPUs and NVIDIA Hopper / Blackwell GPUs in confidential compute mode" (FAQ). The client verifier checks the CPU attestation back to AMD's root certificate (architecture).
- Supported platforms: Tinfoil's primer lists AMD EPYC (SEV-SNP), Intel Xeon 5th Gen / Xeon 6 (TDX), and NVIDIA H100, H200 and B200 GPUs (primer). Its Blackwell benchmark post used an Intel TDX + NVIDIA confidential-computing setup (blog).
- Conflicting claims: some third-party directories describe Tinfoil as "Intel TDX and NVIDIA H100 CC" (confidentialinference.net). Tinfoil's own FAQ says AMD SEV-SNP is what's in production now, so treat the directory as out of date.
- Big models: run on 8 GPUs per enclave. The public config repos for GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash specify
gpus: 8, and the GLM-5.3 and DeepSeek repos say "8x Blackwell" (confidential-glm5-3-nvfp4, confidential-deepseek-v4-1-flash, confidential-kimi-k3). - GPU check: at boot, the enclave checks that each NVIDIA GPU is in confidential-compute mode, and refuses to start if it isn't (architecture).
Attestation and verification, step by step
- Build: Tinfoil's enclave code (firmware, a minimal Ubuntu VM image, the vLLM inference server, the model router) is open source on GitHub. A GitHub Action builds it reproducibly and publishes the expected fingerprints ("measurements") to the Sigstore transparency log (verification, GitHub org).
- Model weights are pinned too. A tool called modelwrap fingerprints the exact Hugging Face weights. The enclave refuses to read any block of weights that doesn't match (dm-verity), so Tinfoil can't silently swap in a cheaper model (architecture, FAQ).
- Connect: on each connection, the SDK fetches the enclave's hardware-signed attestation and the Sigstore bundle. It checks that they match, then pins TLS to the enclave's attested key. If anything fails, it refuses to send data (verification).
- Router chain: you connect to a model router that also runs in an attested enclave. The router verifies each model enclave before forwarding traffic, so the whole chain stays encrypted from the host's point of view (verification, confidential-model-router).
- Body encryption (EHBP): request and response bodies are also encrypted with HPKE to the enclave's attested key. Your own backend proxy (or anything else in between) can route requests but can't read them (EHBP, architecture).
- Audit trail: each enclave boot's attestation is also embedded in a TLS certificate recorded in public Certificate Transparency logs (verification).
What Tinfoil can and can't see
- Can see: network metadata (source IP, packet timing, byte counts) and per-request billing metadata (input token count, output token count, model name, timestamp) (FAQ). Also account, billing, usage and security data (privacy policy).
- Can't see: prompts, completions, uploaded files, embeddings, tool-call payloads, or request and response bodies. "We do not retain prompt or response content after the response is returned, and we do not use API content to train models" (privacy policy, FAQ).
- Legal requests: Tinfoil reports zero government requests as of September 22, 2026. It says it can't hand over content even if compelled (privacy policy).
- Limits Tinfoil admits to: hardware attacks (published key-extraction and attestation-forgery attacks need physical access, see tee.fail), side channels, I/O-pattern leakage, a compromised user device, and denial of service by cloud providers (FAQ, primer).
- Compliance: SOC 2 Type II (privacy policy). US company headquartered in San Francisco (privacy policy).
- Important caveat for Chance AI: the guarantee is only cryptographically checked if you use Tinfoil's SDK or its local proxy. Plain HTTPS "skips connection-time verification" (direct API docs).
- Skeptics: a Reddit thread points out that the data is still plaintext inside the enclave while the model processes it. That's true; the protection comes from hardware isolation, not from encrypting the computation itself (r/LocalLLaMA).
Company, funding, and who uses it
- Founders: Tanya Verma (ex-Cloudflare), Jules Drean (MIT PhD in secure hardware; worked at Microsoft Research and NVIDIA), and Sacha Servan-Schreiber (MIT PhD in cryptography) (company page).
- Batch: Y Combinator, Spring 2025 ("YC P25") (Launch HN, YC).
- Backers named on the company page: Felicis and Y Combinator, plus angels including Paul Graham, Nick Sullivan and Michael Grinich (company page).
- Funding amount: Unverified. Third-party databases disagree: $125K YC seed (Dealroom), "$4M seed (2023)" (StartupHub), and a $7.1M seed with about $7.7M total (LinkedIn data). I found no official announcement.
- Users:
- DuckDuckGo's Duck.ai offers free gpt-oss-120b and Gemma 4 31B "hosted by Tinfoil.sh", labeled "zero provider visibility" (Duck.ai help, Duck.ai).
- Tinfoil's customer page lists Mozilla/MZLA (Thunderbolt), Screenpipe, Workshop Labs, Cybera, UC Berkeley, Harvard researchers and Pour Demain (customers).
- The home page mentions a collaboration with Red Hat (home).
- Products: Private Chat ($20/month), the Private Inference API (usage-based), and Tinfoil Containers ($20/month + usage) (home).
2. Models
Chat models (the full current list)
Sources: chat models docs, official catalog feed. The "AA index" column is the Artificial Analysis Intelligence Index (AA). "Complied" is the share of 1,500 harmful-request test prompts the model answered in Tinfoil's own safety evaluation; higher means less filtered (my calculation from Tinfoil's public data, safeguard-evals).
| Model (API id) | Size | API context | Vision | Reasoning | AA index | Complied |
|---|---|---|---|---|---|---|
Kimi K3 (kimi-k3) |
2.8T total / 104B active | 256K | Yes | Always on (low/high/max) | 44 (max) | 2.7% |
GLM-5.3 (glm-5-3) |
743B / 39B active | 1M | No | Always on (low/high/max) | 45 (max) | 3.8% |
GLM-5.3 Flash (glm-5-3-flash), removed Oct 9 |
320B / 18B active | 1M | Yes | Always on | 42 | 3.4% |
DeepSeek V4.1 Flash (deepseek-v4-1-flash), experimental |
552B MoE | 1M | Yes | Optional (none to max) | 39 (max) / 25 (off) | 1.9% |
Gemma 4 31B (gemma4-31b) |
31B | 256K | Yes | Optional | 15 | 6.6% |
gpt-oss-120b (gpt-oss-120b) |
117B / 5.1B active | 131K | No | low/medium/high | 12 (high) | 0.5% |
Llama 3.3 70B (llama3-3-70b) |
70B | 128K | No | None | 8 | 14.5% |
- Reasoning settings and the GLM "always on, defaults to max" behavior: reasoning guide.
- Kimi K3's context is 256K on Tinfoil, versus 1M at Venice (Venice models API).
- In Tinfoil Chat, GLM and DeepSeek are limited to 524K context (catalog feed).
- Quantization: Tinfoil serves several models at reduced precision. Kimi K3 uses MXFP4, GLM-5.3 uses NVFP4, and DeepSeek uses FP8/FP4 (chat models docs). Because the weights are attested, you can confirm exactly which weights are served (FAQ).
Are any of them low-filter or uncensored like Venice?
- No uncensored or "abliterated" models. Every model is the stock open-weight release (see the weight links in the chat docs). Venice, by contrast, sells Venice Uncensored 1.2, Gemma 4 Uncensored, GLM 4.7 Flash Heretic, Abliterated Large V2 and Qwen 3.6 Plus Uncensored (Venice models API).
- No added moderation layer on the API. "Safeguards apply only to Tinfoil Chat, not the Inference API or Containers" (FAQ). In Chat, a classifier running inside the enclave flags only three "hard-no" areas: child endangerment, mass violence/terrorism, and self-harm. Only a one-bit flag and a conversation ID leave the enclave (safety page, confidential-safeguards). Tinfoil's evaluation repo says it "defer[s] to default model behavior outside a small set of hard-nos" (safeguard-evals).
- Pre-hosting test: every chat model is tested against the hard-no policy before Tinfoil hosts it (safety page).
- The acceptable use policy still applies. It forbids trying to bypass safeguards to get prohibited output (AUP).
- How filtered each model is (built-in behavior). From Tinfoil's published runs of 1,500 harmful prompts (AILuminate + HarmBench, scored by a gpt-4.1-mini judge), here's the share the model answered. Treat this as a rough guide: the judge output format differs slightly between older and newer runs.
- Least filtered: Llama 3.3 70B 14.5%, then Gemma 4 31B 6.6%
- Middle: GLM-5.3 3.8%, GLM-5.3 Flash 3.4%, Kimi K3 2.7%
- Strict: DeepSeek V4.1 Flash 1.9%, gpt-oss-120b 0.5% (very strict)
- Data:
evals/data/rate_<model>.jsonlin safeguard-evals. - Practical meaning: for normal adult topics (frank health, legal or relationship questions, dark fiction), GLM, Kimi and DeepSeek will generally be fine. But none of them is designed to be "uncensored" the way Venice's own models are. Unverified: I didn't run my own refusal tests.
Other model types
- Vision: Kimi K3, GLM-5.3 Flash, DeepSeek V4.1 Flash and Gemma 4 31B (vision docs).
- Speech-to-text:
- Whisper Large V3 Turbo: $0.01 per request
- Voxtral Small 24B (also answers questions about audio): $0.20 in / $0.60 out per 1M
- Voxtral Mini Realtime (streaming over WebSocket; verified connection currently only via the Node JS SDK): $0.002 per request
- Sources: audio docs, catalog feed.
- Text-to-speech: Voxtral TTS ($0.002 per request) and Qwen3 TTS ($0.01 per request) (audio docs, catalog feed).
- Embeddings: nomic-embed-text v1.5 (768 dimensions), $0.05 per 1M (embedding docs, catalog feed).
- Other:
- Document conversion: $0.05 per request
- Web search tool: $0.05 per request
- PII filter: $0.005 per request
- gpt-oss-safeguard-120b classifier: $0.15 / $0.60
- Sources: catalog feed, safety models.
- Image generation: none. No image-generation models are listed in the docs or the catalog (models overview, catalog feed).
3. Intelligence vs Venice
Which Venice models I compared: I pulled Venice's live model list (api.venice.ai/api/v1/models). Venice marks Grok 4.7 with the trait most_intelligent, and its privacy mode is "private." Venice also offers frontier closed models (Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, Gemini) in "anonymized" mode. Venice's own docs say that in that mode "Prompt content is still visible to that provider" (Venice privacy docs). So I compare against both groups.
Side by side (higher is better)
| Model | Where | Privacy on that service | AA Intelligence Index | LMArena text (Oct 2) | LiveBench avg | GPQA Diamond | Humanity's Last Exam | SciCode (coding) | Terminal-Bench 4.0 (agentic coding) |
|---|---|---|---|---|---|---|---|---|---|
| GLM-5.3 (max) | Tinfoil + Venice | TEE (Tinfoil) / private (Venice) | 45 | 1478 | 76.1% | 91.7% | 42.3% | 59.0% | 41.9% |
| Kimi K3 (max) | Tinfoil + Venice | TEE / private | 44 | 1488 | 79.2% | 93.5% | 46.9% | 59.5% | 12.6% |
| DeepSeek V4.1 Flash (max) | Tinfoil + Venice | TEE / private | 39 | 1474 | 81.1% | not published | 39.2% | 51.9% | 26.8% |
| Gemma 4 31B (reasoning) | Tinfoil + Venice | TEE / private | 15 | 1453 | not listed | 85.7% | 23.6% | 45.5% | 0% |
| gpt-oss-120b (high) | Tinfoil + Venice | TEE / private | 12 | 1352 | not listed | 78.2% | 19.6% | 34.0% | 0% |
| Llama 3.3 70B | Tinfoil + Venice | TEE / private | 8 | 1318 | not listed | 49.8% | 3.6% | not published | not published |
| Grok 4.7 (xhigh), Venice's "most intelligent" | Venice only | private (Venice ZDR) | 46 | 1442 | 77.4% | not published | 43.1% | 57.4% | 25.8% |
| Qwen 3.8 2.4T | Venice only | private | 40 | not checked | not checked | not checked | not checked | not checked | not checked |
| DeepSeek V4 Pro 0813 (max) | Venice only | private | 36 | not checked | 77.4% (V4 Pro) | not checked | not checked | not checked | not checked |
| Claude Opus 5.5 (max) | Venice only | anonymized (Anthropic sees prompts) | 58 | 1504 (high) | 83.2% | not published | 61.4% | 66.9% | 59.6% |
| GPT-6 Astra (max) | Venice only | anonymized | 53 | 1477 | 82.2% | 96.1% | 54.7% | 56.5% | 59.1% |
| Claude Fable 5.1 (max) | Venice only | anonymized | 53 | 1501 | 83.4% | 93.7% | 59.1% | 63.1% | 52.0% |
| Reference: Gemini 3.8 Flash (high) | OpenRouter | n/a | 41 | 1495 | not checked | 95.3% | 47.8% | 56.6% | 19.7% |
Where the numbers come from:
- AA Intelligence Index, GPQA, HLE, SciCode, Terminal-Bench 4.0: Artificial Analysis, read on Oct 6, 2026 (leaderboard, GPQA page, and the model pages for Gemma 4 31B, gpt-oss-120b, Llama 3.3 70B, Grok 4.7 and DeepSeek V4.1 Flash). AA no longer publishes MMLU-Pro for these models; its current index (v4.3.2) uses newer tests (AA index).
- LMArena text Elo: Oct 2, 2026 snapshot (arena.ai/leaderboard/text).
- LiveBench: as mirrored by BenchLeader, data as of Oct 5, 2026 (BenchLeader LiveBench; original at livebench.ai).
- Caveat: these are scores for each model in general, not for Tinfoil's servers specifically. Tinfoil runs some quantized versions (see Models), and no independent group benchmarks Tinfoil-served quality. Unverified.
Plain answer
- Does Tinfoil have a model as smart as Venice's best?
- Against Venice's private models: yes, essentially. GLM-5.3 (45) is one point behind Grok 4.7 (46) on Artificial Analysis. Kimi K3 beats Grok 4.7 on LMArena (1488 vs 1442), and DeepSeek V4.1 Flash beats it on LiveBench (81.1% vs 77.4%). Tinfoil also hosts the very same GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash models Venice runs privately (Venice models API).
- Against the frontier models Venice resells (Claude Opus 5.5 at 58, GPT-6 Astra and Claude Fable 5.1 at 53): no. Tinfoil is roughly 8–13 index points behind. On Venice those models are only "anonymized," so they aren't private in the way Tinfoil is.
- Smart slot: GLM-5.3 (
glm-5-3). It has the highest AA index on Tinfoil, the best agentic-coding score, a 1M context, and costs under a third of Kimi K3's output price ($5.75 vs $20). - Use Kimi K3 as a "Max" option for the hardest prompts or image analysis, if the cost is acceptable.
4. Pricing and plans
API prices per 1M tokens (official catalog feed)
| Model | Input | Cached input | Output |
|---|---|---|---|
| gpt-oss-120b | $0.15 | n/a | $0.60 |
| Gemma 4 31B | $0.40 | n/a | $1.00 |
| GLM-5.3 Flash (removed Oct 9) | $0.40 | $0.10 | $1.25 |
| DeepSeek V4.1 Flash | $0.65 | $0.13 | $1.45 |
| Llama 3.3 70B | $1.75 | n/a | $2.75 |
| GLM-5.3 | $1.80 | $0.45 | $5.75 |
| Kimi K3 | $4.00 | $0.80 | $20.00 |
Sources: api.tinfoil.sh/api/config/models. The same numbers appear on endpoints.run.
How billing works
- API is pay-as-you-go: "API access is pay-as-you-go with per-model pricing based on input and output token usage" (pricing).
- Starter credit: activating Private Inference creates a default key and "adds $1 in API credits to your account" (get an API key). A 2025 Product Hunt note mentioned $5 in credits; that's outdated (chatgate summary).
- Card required: you add a card at activation and are "only charged based on usage." The setup flow calls this "activat[ing] your API subscription," but the charges are usage-based (get an API key).
- Billing period: usage is tracked per "token billing period," and invoices come through Stripe (admin API).
- Prepaid credits: the pricing page also mentions "credit purchases" (pricing).
- Minimum top-up amount: Unverified. It isn't published publicly; it may be visible inside the dashboard.
- Spending controls: per-key spend limits and token caps, plus an account-wide usage limit that pauses the API (get an API key). When you hit a cap, the API returns HTTP 429/402
insufficient_quota(error docs). - Payment methods: "All payments are handled through Stripe for our self-serve plans" (pricing). Mobile Chat subscriptions go through Apple or Google via RevenueCat (privacy policy).
- Crypto: none found. No crypto option appears on the pricing page, terms or docs (Unverified, absence only). For comparison, Venice and PPQ.AI accept crypto (Venice pricing).
- Subscriptions: Private Chat is $20/month (web + iOS) (pricing). It's separate from the API, and you don't need it for Chance AI. Containers are $20/month + usage (home).
- Price versus non-private hosts: Tinfoil charges a premium. For example, GLM-5.3 is $1.40 / $4.40 at Z.AI on OpenRouter, Kimi K3 is $3 / $15 at Moonshot, and DeepSeek V4.1 Flash is $0.15 / $0.60 at DeepSeek (OpenRouter GLM-5.3, Kimi K3, DeepSeek V4.1 Flash). Venice charges $1.75 / $5.50 for GLM-5.3 and $0.30 / $1.20 for DeepSeek V4.1 Flash (Venice models API).
5. API
The basics
- Base URL:
https://inference.tinfoil.sh/v1, OpenAI-compatible (direct API docs). - Auth:
Authorization: Bearer <TINFOIL_API_KEY>(direct API docs). - Python SDK:
pip install tinfoil, thenfrom tinfoil import TinfoilAI. It's a wrapper around the official OpenAI client: "All method signatures and response formats remain the same." It also has an async client,AsyncTinfoilAI(Python SDK). Latest PyPI version is 0.14.0, which requires Python 3.10 or newer (PyPI).
Plain openai or requests without the SDK
This works: point OpenAI(base_url="https://inference.tinfoil.sh/v1") at it. What you lose:
- Verification. You get no proof that you're talking to an enclave, and no protection against man-in-the-middle attacks. Tinfoil literally labels this "No privacy guarantee" and says it's "not recommended for production" (direct API docs).
- Speed. Without the SDK you lose its router and enclave load balancing, which "can degrade performance" (direct API docs).
- Cache separation. The SDK automatically adds a prompt-cache secret; without it you'd need to set one yourself (prompt caching).
Middle option: run the tinfoil-proxy Docker container next to the app. It verifies attestation itself and exposes a plain OpenAI-style endpoint at http://127.0.0.1:3301/v1 (proxy CLI).
Features
- Streaming: yes, with
stream=True(Python SDK). - Tool/function calling: yes, standard
tools/tool_choice. Most chat models support it, and Kimi K3 is recommended for complex tool use (tool calling, catalog feed). - JSON / structured outputs: yes,
response_formatwithjson_schemavia vLLM; there are alsochoiceandregexconstraints throughextra_body(structured outputs). Unverified: I found no mention of the simplerjson_objectmode. - Reasoning control:
reasoning_effort. Kimi K3 and GLM-5.3 always reason and default to "max", and reasoning tokens count towardmax_tokens, so setreasoning_effort="low"for everyday chat (reasoning guide). - Other features: prompt caching (automatic in SDK 0.13.5 and later), images as base64, and a
/v1/responsesendpoint on some models (prompt caching, catalog feed).
Flask sketch: Tinfoil as a second provider next to OpenRouter
# providers.py (keys come from environment variables, never hard-coded)
import os
from openai import OpenAI
from tinfoil import TinfoilAI # pip install tinfoil (Python 3.10+)
openrouter = OpenAI(base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"])
tinfoil = TinfoilAI(api_key=os.environ["TINFOIL_API_KEY"]) # verifies the enclave before sending data
# Picker slot -> model id, per backend
MODELS = {
"openrouter": {"default": "google/gemini-3.8-flash", "smart": "google/gemini-3.1-pro-preview"},
"tinfoil": {"default": "deepseek-v4-1-flash", "smart": "glm-5-3",
"max": "kimi-k3", "vision": "gemma4-31b", "loose": "llama3-3-70b"},
}
REASONING = {"glm-5-3": "low", "kimi-k3": "low", "deepseek-v4-1-flash": "low"} # avoid default "max"
def stream_chat(messages, slot="default", private=False):
backend = "tinfoil" if private else "openrouter"
client = tinfoil if private else openrouter
model = MODELS[backend].get(slot, MODELS[backend]["default"])
kwargs = {"reasoning_effort": REASONING[model]} if private and model in REASONING else {}
stream = client.chat.completions.create(model=model, messages=messages,
stream=True, max_tokens=8000, **kwargs)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
yield chunk.choices[0].delta.content
# app.py
from flask import Flask, Response, request
from providers import stream_chat
app = Flask(__name__)
@app.post("/api/chat")
def chat():
d = request.get_json()
gen = stream_chat(d["messages"], d.get("slot", "default"), d.get("private", False))
return Response(gen, mimetype="text/plain")
Notes:
- The model ids,
reasoning_effortvalues and streaming pattern come from Tinfoil's docs (chat models, reasoning guide, Python SDK). The OpenRouter model ids are only examples. - Create the clients once at start-up and reuse them.
- Keep the model map in config: Tinfoil retires models often (see Downsides).
- If you ever add many users, pass a distinct
user_cache_secretper user (prompt caching). - I didn't test this code against the live API, since that would have meant signing up.
6. Rate limits, reliability, uptime
Rate limits
- Limits are per model and per usage tier, shared across your keys, in requests per minute and tokens per minute (pricing).
- Your tier rises automatically with lifetime paid usage; promotional credits don't count (pricing).
- Actual numbers: Unverified. They're shown only in the dashboard's Limits tab.
Uptime promises
Tinfoil "target[s] 99.9% monthly uptime" but does "not guarantee uptime above 95%." There are no SLA credits unless you have an enterprise contract (terms).
Status page
At status.tinfoil.sh, read Oct 6, 2026 at 4:10 AM PT, all services were online. Per-model uptime over the history the page shows (its bars span about 180 days):
| Component | Uptime |
|---|---|
| gpt-oss-120b | 99.935% |
| llama3-3-70b | 99.914% |
| gemma4-31b | 99.630% |
| kimi-k3 | 99.473% |
| glm-5-3 | 99.466% |
| glm-5-3-flash | 98.435% |
| deepseek-v4-1-flash | 97.873% |
| Backend infrastructure | 99.986% |
| KDS attestation proxy | 98.781% |
Recent incidents (status RSS; times converted to PT)
- DeepSeek V4.1 Flash went down four times recently:
- Sep 26, 1:11–3:32 AM
- Oct 1, 6:09–7:42 PM
- Oct 4, 2:49–4:23 PM
- plus a Sep 20 outage
- Kimi K3 was down Sep 25, 2:22–3:02 PM.
- Multi-service outage Sep 21 starting 11:30 AM: several model endpoints were down about 15 minutes, and "Chat Models" took until about 6:03 PM to recover.
- Many short "Chat Models" outages Sep 13–22.
- Earlier outage: in April 2026, models (gpt-oss-120b, Gemma and others) were unreachable for about 30–60 minutes (status page).
Speed vs OpenRouter
- Unverified: no independent Tinfoil-specific speed data.
- Artificial Analysis doesn't list Tinfoil as a provider for these models (AA Kimi K3 providers).
- Tinfoil isn't on OpenRouter (OpenRouter providers).
- endpoints.run shows no speed or latency figures for it (endpoints.run).
- Cross-provider medians from AA (as a rough guide): DeepSeek V4.1 Flash runs about 222–311 tokens/second, GLM-5.3 about 75–83, Kimi K3 about 52–55, versus about 243 for Gemini 3.8 Flash (AA leaderboard).
- Tinfoil's own benchmarks show confidential computing adds roughly 1% to about 30% to the time between tokens, and over 100% at low concurrency, depending on the inference engine (CC overhead blog).
- Expect: Tinfoil to be somewhat slower than the fastest OpenRouter providers, and noticeably slower on Kimi K3.
7. Comparison table
| Tinfoil | Venice | Privatemode | Scaleway Generative APIs | |
|---|---|---|---|---|
| Privacy guarantee | Hardware-verified: TEE plus client-side attestation, open source (docs) | Mostly policy. Four modes: Anonymous (provider sees prompts), Private (contract ZDR), TEE, and E2EE on e2ee-* models (Venice privacy) |
Hardware-verified: confidential computing plus end-to-end encryption via its proxy (Privatemode) | Policy: zero data retention by default (Scaleway privacy) |
| Jurisdiction | USA, San Francisco (privacy) | USA, Wyoming law (Venice ToS) | Germany, Edgeless Systems GmbH, Bochum (imprint) | France/EU, Paris; says it's "not subject to … American Cloud Act" (Scaleway privacy) |
| Models | 7 chat models (6 after Oct 9) plus audio, embeddings, vision; open-weight only (chat docs) | 100+ text models, including Grok, Claude, GPT and Gemini (anonymized) plus many open models (Venice models API) | 3 chat models (GLM-5.3, GLM-5.3 Flash, gpt-oss-120b) plus OCR (pricing) | Around 9+ open models (gpt-oss, Gemma, Qwen, DeepSeek, Llama, Mistral, GLM-5.2) (pricing) |
| Filtering | Stock model behavior; no API moderation (FAQ) | Lowest. Dedicated uncensored models (Venice models API) | Stock model behavior (gpt-oss is strict). Unverified whether any added moderation exists | Stock model behavior. Unverified whether any added moderation exists |
| Pricing model | Pay-as-you-go per token, $1 starter credit (pricing, key docs) | Pay-as-you-go credits, or Pro $18/month and higher; crypto accepted (Venice pricing) | Free tier (5M tokens at sign-up), then pay-as-you-go in EUR (pricing) | Pay-as-you-go per token, 1M free tokens (pricing) |
| OpenAI-compatible | Yes. SDK is a drop-in; raw HTTPS works but is unverified (Python SDK) | Yes | Yes, through its encryption proxy (Privatemode) | Yes |
| Ease of setup | Easy: pip install tinfoil, change one line |
Easiest: plain OpenAI client, or via OpenRouter | Medium: run a proxy container | Easy: plain OpenAI client |
| On OpenRouter? | No (OpenRouter providers) | Yes (OpenRouter providers) | No | No |
The Privatemode and Scaleway rows rely partly on the earlier research file from today (privacy-llm-providers-2026-10-06.md). I re-checked Privatemode's prices and imprint, Venice's model list and terms, and OpenRouter's provider list live today.
8. Downsides and limitations
- No Gemini, Claude, Grok or GPT-5/6. Only open-weight models are hosted (models overview). Chance AI's Gemini, Claude and Grok picker slots have no Tinfoil equivalent.
- Small menu that changes often. Seven chat models today, six after GLM-5.3 Flash is removed on Oct 9. Kimi K2.6 and DeepSeek V4 Pro were removed in July 2026, and GLM-5.2 and DeepSeek V4 Flash in September (changelog). Model ids will break, so keep them in config.
- No uncensored models. That's a real gap compared with Venice for Justin's low-filter preference. gpt-oss-120b is very strict (safeguard-evals data).
- Price premium over non-private hosts: about 1.3x for GLM-5.3 and Kimi K3, about 4x on input for DeepSeek V4.1 Flash, versus the cheapest first-party prices on OpenRouter (OpenRouter endpoints). Kimi K3 is expensive at $20 per 1M output (catalog feed).
- Hidden reasoning cost. GLM-5.3 and Kimi K3 always reason and default to "max," and reasoning tokens are billed as output (reasoning guide).
- Speed: confidential-computing overhead, plus Kimi K3 being slow (see section 6). No independent Tinfoil-specific measurements exist.
- Maturity: a young company (YC Spring 2025). Weeks of short outages in Sept 2026. The only uptime guarantee is 95% (terms, status RSS).
- Lock-in is low but real. The API is OpenAI-compatible, but keeping the verified guarantee means depending on Tinfoil's SDK (Python 3.10+) or its proxy (PyPI, proxy CLI).
- Privacy is strong, not absolute. Network metadata is visible, and hardware or side-channel attacks remain possible (FAQ). Chance AI's own server still sees plaintext before it sends to Tinfoil, so the app itself has to be trustworthy.
- No crypto payments, and minimum top-up is unpublished (Unverified; see Pricing).
9. Verdict
Yes: Tinfoil suits a "Private mode" backend toggle next to OpenRouter in Chance AI. It is the only option here that is hardware-verified, pay-as-you-go, drop-in for Python, and independent of OpenRouter.
Recommended picker mapping (when Private mode is on)
| Chance AI slot | OpenRouter today | Tinfoil (private) | Notes |
|---|---|---|---|
| Default (Gemini Flash) | Gemini Flash | DeepSeek V4.1 Flash | AA 39 vs 41 for Gemini 3.8 Flash. Set reasoning_effort="low". Fall back to GLM-5.3 if it's down (it's "experimental") |
| Pro / smart (Gemini Pro, Claude) | Gemini Pro / Claude | GLM-5.3 | Smartest Tinfoil model. reasoning_effort="low" normally, "high" for hard tasks |
| Grok / DeepSeek | Grok / DeepSeek | Kimi K3 ("Max") | Hardest questions and images; about 2.4x GLM-5.3's monthly cost at the same usage |
| Llama | Llama | Llama 3.3 70B | Least filtered Tinfoil model, but weak |
| Mistral / cheap vision | Mistral | Gemma 4 31B | Cheap, reads images |
| Claude, Gemini, Grok as the real models | n/a | None | Disable those slots, or show "not available in Private mode" |
For low-filter chats, keep a Venice toggle (directly, or through OpenRouter's Venice endpoints). Tinfoil won't replace it.
Estimated monthly cost: one heavy user (1,200 messages × 6K tokens in, 400 out)
That's 7.2M input tokens and 0.48M output tokens a month. Prices come from the catalog feed. These are my calculations.
| Setup | Monthly cost |
|---|---|
| All DeepSeek V4.1 Flash | $5.38 (about $3.50 if half the input hits the prompt cache) |
| All GLM-5.3 | $15.72 (about $10.86 with 50% cache hits) |
| All Kimi K3 | $38.40 |
| All gpt-oss-120b (cheapest, strict) | $1.37 |
| Recommended mix: 70% DeepSeek V4.1 Flash, 20% GLM-5.3, 10% Kimi K3 | ≈ $10.75 base. Range about $7–$14: about $7.31 with 50% cache hits, about $13.75 if each answer adds roughly 600 hidden reasoning tokens |
Cost assumptions (Unverified): the 50% cache-hit rate and 600 reasoning tokens are illustrative guesses, not measurements. Real numbers depend on how much chat history repeats and on the reasoning effort setting. For comparison, the same 1,200 messages on Venice's Grok 4.7 would cost about $19.61 ($2.27 / $6.80 per 1M, Venice models API).
Practical tip: set a monthly spend limit (for example $20) on the Tinfoil key so the toggle can never surprise you (get an API key).
10. Sources (all accessed Oct 6, 2026)
Tinfoil, official: - tinfoil.sh - Pricing - Model catalog, with its feed at api.tinfoil.sh/api/config/models - Private Inference - Privacy policy (Sept 22, 2026) - Security & Privacy FAQ (Sept 25, 2026) - Terms - Acceptable use policy - Safety & Safeguards - Company - Customers - Technology - Blog, especially the confidential-computing overhead post - Status page and status RSS
Tinfoil docs: - Docs index (llms.txt) - Models: Chat · Vision · Audio · Embeddings · Safety · Overview - Getting started: API key · Python SDK · Direct API · Proxy CLI - Guides: Reasoning · Tool calling · Structured outputs · Prompt caching · Errors · Admin API - Security: Verification · Architecture · Enclave primer · EHBP - Changelog
Tinfoil code: - github.com/tinfoilsh - safeguard-evals · confidential-safeguards · confidential-model-router - Model configs: Kimi K3 · GLM-5.3 · DeepSeek V4.1 Flash - PyPI: tinfoil
Community and press: - Launch HN - r/LocalLLaMA thread - Y Combinator profile - Funding databases: Dealroom · StartupHub · LinkedIn - chatgate (Product Hunt summary) - Duck.ai: privacy help page · Duck.ai help
Price and directory trackers: - endpoints.run/providers/tinfoil - confidentialinference.net - llmprice
Benchmarks: - Artificial Analysis: leaderboard · Intelligence Index · GPQA · HLE · SciCode - AA model pages: Gemma 4 31B · gpt-oss-120b · Llama 3.3 70B · Grok 4.7 · DeepSeek V4.1 Flash - LMArena text - LiveBench: via BenchLeader · livebench.ai
Venice: - Models API - Privacy docs - Pricing - Terms of service
OpenRouter: - Providers list - Endpoints: GLM-5.3 · Kimi K3 · DeepSeek V4.1 Flash
Privatemode: - Pricing - Imprint - Inference API
Scaleway: - Data privacy - Pricing
Prior research: privacy-llm-providers-2026-10-06.md, Chance AI research folder, Oct 6, 2026.