Tinfoil (tinfoil.sh): Deep-Dive Report for Chance AI

Privacy-focused AI inference: how it works, models, intelligence vs Venice, pricing, API, uptime, verdict. October 6, 2026.

Researched Tuesday, October 6, 2026 (PT). Every claim links to its source. Anything I could not confirm is marked Unverified. I didn't sign up for anything or spend money. Prices are in US dollars per 1 million tokens ("per 1M") unless noted.

Bottom line

Top 5 Tinfoil models for Chance AI (ranked)

Ranked on four things: intelligence compared with Venice, filtering, speed, and price. Prices come from Tinfoil's official model catalog feed (api.tinfoil.sh/api/config/models, which powers tinfoil.sh/models).

# Model (API id) Best for Price in / out per 1M
1 GLM-5.3 (glm-5-3) "Smart" slot. The smartest Tinfoil model (AA index 45), close to Venice's best private model, with a 1M-token context. $1.80 / $5.75 (cached input $0.45)
2 DeepSeek V4.1 Flash (deepseek-v4-1-flash) Everyday default. Nearly as smart as Gemini 3.8 Flash (AA 39 vs 41), very fast, reads images, 1M context. Marked "experimental" and has had the most outages lately. $0.65 / $1.45 (cached $0.13)
3 Kimi K3 (kimi-k3) "Max" slot for the hardest questions and image analysis. Highest LMArena score of any Tinfoil model, but slow and expensive. $4.00 / $20.00 (cached $0.80)
4 Gemma 4 31B (gemma4-31b) Cheap image reading and light chat. The second-least filtered Tinfoil model in Tinfoil's own safety tests. $0.40 / $1.00
5 Llama 3.3 70B (llama3-3-70b) "Looser" slot. The least filtered Tinfoil model by a wide margin, but much less smart (AA 8) and pricey for its quality. $1.75 / $2.75

Not in the top 5: gpt-oss-120b is the cheapest at $0.15 / $0.60, but it's the most restrictive model Tinfoil hosts (it complied with only 0.5% of test prompts) and much less smart (AA 12). GLM-5.3 Flash is deprecated and will be removed October 9, 2026 (changelog).

Contents

  1. What Tinfoil is and how it works
  2. Models
  3. Intelligence vs Venice
  4. Pricing and plans
  5. API and a Flask code sketch
  6. Rate limits, reliability, uptime
  7. Comparison: Tinfoil vs Venice vs Privatemode vs Scaleway
  8. Downsides and limitations
  9. Verdict and monthly cost
  10. Sources

1. What Tinfoil is and how it works

In plain English

Normally an AI provider decrypts your message on its servers, and its staff, its cloud host, or a court order could in principle get at it. Tinfoil runs each model inside a confidential virtual machine. The chip encrypts that machine's memory so the server operator (Tinfoil, or its cloud provider) can't look inside. Before your app sends anything, Tinfoil's SDK asks the chip for a signed "fingerprint" of exactly what software and model weights are running. It compares that fingerprint with the one Tinfoil published publicly from its open-source code. Only if they match does it encrypt your message to a key that exists only inside that enclave (verification docs, enclave primer).

Hardware

Attestation and verification, step by step

  1. Build: Tinfoil's enclave code (firmware, a minimal Ubuntu VM image, the vLLM inference server, the model router) is open source on GitHub. A GitHub Action builds it reproducibly and publishes the expected fingerprints ("measurements") to the Sigstore transparency log (verification, GitHub org).
  2. Model weights are pinned too. A tool called modelwrap fingerprints the exact Hugging Face weights. The enclave refuses to read any block of weights that doesn't match (dm-verity), so Tinfoil can't silently swap in a cheaper model (architecture, FAQ).
  3. Connect: on each connection, the SDK fetches the enclave's hardware-signed attestation and the Sigstore bundle. It checks that they match, then pins TLS to the enclave's attested key. If anything fails, it refuses to send data (verification).
  4. Router chain: you connect to a model router that also runs in an attested enclave. The router verifies each model enclave before forwarding traffic, so the whole chain stays encrypted from the host's point of view (verification, confidential-model-router).
  5. Body encryption (EHBP): request and response bodies are also encrypted with HPKE to the enclave's attested key. Your own backend proxy (or anything else in between) can route requests but can't read them (EHBP, architecture).
  6. Audit trail: each enclave boot's attestation is also embedded in a TLS certificate recorded in public Certificate Transparency logs (verification).

What Tinfoil can and can't see

Company, funding, and who uses it

2. Models

Chat models (the full current list)

Sources: chat models docs, official catalog feed. The "AA index" column is the Artificial Analysis Intelligence Index (AA). "Complied" is the share of 1,500 harmful-request test prompts the model answered in Tinfoil's own safety evaluation; higher means less filtered (my calculation from Tinfoil's public data, safeguard-evals).

Model (API id) Size API context Vision Reasoning AA index Complied
Kimi K3 (kimi-k3) 2.8T total / 104B active 256K Yes Always on (low/high/max) 44 (max) 2.7%
GLM-5.3 (glm-5-3) 743B / 39B active 1M No Always on (low/high/max) 45 (max) 3.8%
GLM-5.3 Flash (glm-5-3-flash), removed Oct 9 320B / 18B active 1M Yes Always on 42 3.4%
DeepSeek V4.1 Flash (deepseek-v4-1-flash), experimental 552B MoE 1M Yes Optional (none to max) 39 (max) / 25 (off) 1.9%
Gemma 4 31B (gemma4-31b) 31B 256K Yes Optional 15 6.6%
gpt-oss-120b (gpt-oss-120b) 117B / 5.1B active 131K No low/medium/high 12 (high) 0.5%
Llama 3.3 70B (llama3-3-70b) 70B 128K No None 8 14.5%

Are any of them low-filter or uncensored like Venice?

Other model types

3. Intelligence vs Venice

Which Venice models I compared: I pulled Venice's live model list (api.venice.ai/api/v1/models). Venice marks Grok 4.7 with the trait most_intelligent, and its privacy mode is "private." Venice also offers frontier closed models (Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, Gemini) in "anonymized" mode. Venice's own docs say that in that mode "Prompt content is still visible to that provider" (Venice privacy docs). So I compare against both groups.

Side by side (higher is better)

Model Where Privacy on that service AA Intelligence Index LMArena text (Oct 2) LiveBench avg GPQA Diamond Humanity's Last Exam SciCode (coding) Terminal-Bench 4.0 (agentic coding)
GLM-5.3 (max) Tinfoil + Venice TEE (Tinfoil) / private (Venice) 45 1478 76.1% 91.7% 42.3% 59.0% 41.9%
Kimi K3 (max) Tinfoil + Venice TEE / private 44 1488 79.2% 93.5% 46.9% 59.5% 12.6%
DeepSeek V4.1 Flash (max) Tinfoil + Venice TEE / private 39 1474 81.1% not published 39.2% 51.9% 26.8%
Gemma 4 31B (reasoning) Tinfoil + Venice TEE / private 15 1453 not listed 85.7% 23.6% 45.5% 0%
gpt-oss-120b (high) Tinfoil + Venice TEE / private 12 1352 not listed 78.2% 19.6% 34.0% 0%
Llama 3.3 70B Tinfoil + Venice TEE / private 8 1318 not listed 49.8% 3.6% not published not published
Grok 4.7 (xhigh), Venice's "most intelligent" Venice only private (Venice ZDR) 46 1442 77.4% not published 43.1% 57.4% 25.8%
Qwen 3.8 2.4T Venice only private 40 not checked not checked not checked not checked not checked not checked
DeepSeek V4 Pro 0813 (max) Venice only private 36 not checked 77.4% (V4 Pro) not checked not checked not checked not checked
Claude Opus 5.5 (max) Venice only anonymized (Anthropic sees prompts) 58 1504 (high) 83.2% not published 61.4% 66.9% 59.6%
GPT-6 Astra (max) Venice only anonymized 53 1477 82.2% 96.1% 54.7% 56.5% 59.1%
Claude Fable 5.1 (max) Venice only anonymized 53 1501 83.4% 93.7% 59.1% 63.1% 52.0%
Reference: Gemini 3.8 Flash (high) OpenRouter n/a 41 1495 not checked 95.3% 47.8% 56.6% 19.7%

Where the numbers come from:

Plain answer

4. Pricing and plans

API prices per 1M tokens (official catalog feed)

Model Input Cached input Output
gpt-oss-120b $0.15 n/a $0.60
Gemma 4 31B $0.40 n/a $1.00
GLM-5.3 Flash (removed Oct 9) $0.40 $0.10 $1.25
DeepSeek V4.1 Flash $0.65 $0.13 $1.45
Llama 3.3 70B $1.75 n/a $2.75
GLM-5.3 $1.80 $0.45 $5.75
Kimi K3 $4.00 $0.80 $20.00

Sources: api.tinfoil.sh/api/config/models. The same numbers appear on endpoints.run.

How billing works

5. API

The basics

Plain openai or requests without the SDK

This works: point OpenAI(base_url="https://inference.tinfoil.sh/v1") at it. What you lose:

  1. Verification. You get no proof that you're talking to an enclave, and no protection against man-in-the-middle attacks. Tinfoil literally labels this "No privacy guarantee" and says it's "not recommended for production" (direct API docs).
  2. Speed. Without the SDK you lose its router and enclave load balancing, which "can degrade performance" (direct API docs).
  3. Cache separation. The SDK automatically adds a prompt-cache secret; without it you'd need to set one yourself (prompt caching).

Middle option: run the tinfoil-proxy Docker container next to the app. It verifies attestation itself and exposes a plain OpenAI-style endpoint at http://127.0.0.1:3301/v1 (proxy CLI).

Features

Flask sketch: Tinfoil as a second provider next to OpenRouter

# providers.py  (keys come from environment variables, never hard-coded)
import os
from openai import OpenAI
from tinfoil import TinfoilAI            # pip install tinfoil  (Python 3.10+)

openrouter = OpenAI(base_url="https://openrouter.ai/api/v1",
                    api_key=os.environ["OPENROUTER_API_KEY"])
tinfoil = TinfoilAI(api_key=os.environ["TINFOIL_API_KEY"])   # verifies the enclave before sending data

# Picker slot -> model id, per backend
MODELS = {
    "openrouter": {"default": "google/gemini-3.8-flash", "smart": "google/gemini-3.1-pro-preview"},
    "tinfoil":    {"default": "deepseek-v4-1-flash", "smart": "glm-5-3",
                   "max": "kimi-k3", "vision": "gemma4-31b", "loose": "llama3-3-70b"},
}
REASONING = {"glm-5-3": "low", "kimi-k3": "low", "deepseek-v4-1-flash": "low"}  # avoid default "max"

def stream_chat(messages, slot="default", private=False):
    backend = "tinfoil" if private else "openrouter"
    client = tinfoil if private else openrouter
    model = MODELS[backend].get(slot, MODELS[backend]["default"])
    kwargs = {"reasoning_effort": REASONING[model]} if private and model in REASONING else {}
    stream = client.chat.completions.create(model=model, messages=messages,
                                            stream=True, max_tokens=8000, **kwargs)
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            yield chunk.choices[0].delta.content

# app.py
from flask import Flask, Response, request
from providers import stream_chat
app = Flask(__name__)

@app.post("/api/chat")
def chat():
    d = request.get_json()
    gen = stream_chat(d["messages"], d.get("slot", "default"), d.get("private", False))
    return Response(gen, mimetype="text/plain")

Notes:

6. Rate limits, reliability, uptime

Rate limits

Uptime promises

Tinfoil "target[s] 99.9% monthly uptime" but does "not guarantee uptime above 95%." There are no SLA credits unless you have an enterprise contract (terms).

Status page

At status.tinfoil.sh, read Oct 6, 2026 at 4:10 AM PT, all services were online. Per-model uptime over the history the page shows (its bars span about 180 days):

Component Uptime
gpt-oss-120b 99.935%
llama3-3-70b 99.914%
gemma4-31b 99.630%
kimi-k3 99.473%
glm-5-3 99.466%
glm-5-3-flash 98.435%
deepseek-v4-1-flash 97.873%
Backend infrastructure 99.986%
KDS attestation proxy 98.781%

Recent incidents (status RSS; times converted to PT)

Speed vs OpenRouter

7. Comparison table

Tinfoil Venice Privatemode Scaleway Generative APIs
Privacy guarantee Hardware-verified: TEE plus client-side attestation, open source (docs) Mostly policy. Four modes: Anonymous (provider sees prompts), Private (contract ZDR), TEE, and E2EE on e2ee-* models (Venice privacy) Hardware-verified: confidential computing plus end-to-end encryption via its proxy (Privatemode) Policy: zero data retention by default (Scaleway privacy)
Jurisdiction USA, San Francisco (privacy) USA, Wyoming law (Venice ToS) Germany, Edgeless Systems GmbH, Bochum (imprint) France/EU, Paris; says it's "not subject to … American Cloud Act" (Scaleway privacy)
Models 7 chat models (6 after Oct 9) plus audio, embeddings, vision; open-weight only (chat docs) 100+ text models, including Grok, Claude, GPT and Gemini (anonymized) plus many open models (Venice models API) 3 chat models (GLM-5.3, GLM-5.3 Flash, gpt-oss-120b) plus OCR (pricing) Around 9+ open models (gpt-oss, Gemma, Qwen, DeepSeek, Llama, Mistral, GLM-5.2) (pricing)
Filtering Stock model behavior; no API moderation (FAQ) Lowest. Dedicated uncensored models (Venice models API) Stock model behavior (gpt-oss is strict). Unverified whether any added moderation exists Stock model behavior. Unverified whether any added moderation exists
Pricing model Pay-as-you-go per token, $1 starter credit (pricing, key docs) Pay-as-you-go credits, or Pro $18/month and higher; crypto accepted (Venice pricing) Free tier (5M tokens at sign-up), then pay-as-you-go in EUR (pricing) Pay-as-you-go per token, 1M free tokens (pricing)
OpenAI-compatible Yes. SDK is a drop-in; raw HTTPS works but is unverified (Python SDK) Yes Yes, through its encryption proxy (Privatemode) Yes
Ease of setup Easy: pip install tinfoil, change one line Easiest: plain OpenAI client, or via OpenRouter Medium: run a proxy container Easy: plain OpenAI client
On OpenRouter? No (OpenRouter providers) Yes (OpenRouter providers) No No

The Privatemode and Scaleway rows rely partly on the earlier research file from today (privacy-llm-providers-2026-10-06.md). I re-checked Privatemode's prices and imprint, Venice's model list and terms, and OpenRouter's provider list live today.

8. Downsides and limitations

  1. No Gemini, Claude, Grok or GPT-5/6. Only open-weight models are hosted (models overview). Chance AI's Gemini, Claude and Grok picker slots have no Tinfoil equivalent.
  2. Small menu that changes often. Seven chat models today, six after GLM-5.3 Flash is removed on Oct 9. Kimi K2.6 and DeepSeek V4 Pro were removed in July 2026, and GLM-5.2 and DeepSeek V4 Flash in September (changelog). Model ids will break, so keep them in config.
  3. No uncensored models. That's a real gap compared with Venice for Justin's low-filter preference. gpt-oss-120b is very strict (safeguard-evals data).
  4. Price premium over non-private hosts: about 1.3x for GLM-5.3 and Kimi K3, about 4x on input for DeepSeek V4.1 Flash, versus the cheapest first-party prices on OpenRouter (OpenRouter endpoints). Kimi K3 is expensive at $20 per 1M output (catalog feed).
  5. Hidden reasoning cost. GLM-5.3 and Kimi K3 always reason and default to "max," and reasoning tokens are billed as output (reasoning guide).
  6. Speed: confidential-computing overhead, plus Kimi K3 being slow (see section 6). No independent Tinfoil-specific measurements exist.
  7. Maturity: a young company (YC Spring 2025). Weeks of short outages in Sept 2026. The only uptime guarantee is 95% (terms, status RSS).
  8. Lock-in is low but real. The API is OpenAI-compatible, but keeping the verified guarantee means depending on Tinfoil's SDK (Python 3.10+) or its proxy (PyPI, proxy CLI).
  9. Privacy is strong, not absolute. Network metadata is visible, and hardware or side-channel attacks remain possible (FAQ). Chance AI's own server still sees plaintext before it sends to Tinfoil, so the app itself has to be trustworthy.
  10. No crypto payments, and minimum top-up is unpublished (Unverified; see Pricing).

9. Verdict

Yes: Tinfoil suits a "Private mode" backend toggle next to OpenRouter in Chance AI. It is the only option here that is hardware-verified, pay-as-you-go, drop-in for Python, and independent of OpenRouter.

Recommended picker mapping (when Private mode is on)

Chance AI slot OpenRouter today Tinfoil (private) Notes
Default (Gemini Flash) Gemini Flash DeepSeek V4.1 Flash AA 39 vs 41 for Gemini 3.8 Flash. Set reasoning_effort="low". Fall back to GLM-5.3 if it's down (it's "experimental")
Pro / smart (Gemini Pro, Claude) Gemini Pro / Claude GLM-5.3 Smartest Tinfoil model. reasoning_effort="low" normally, "high" for hard tasks
Grok / DeepSeek Grok / DeepSeek Kimi K3 ("Max") Hardest questions and images; about 2.4x GLM-5.3's monthly cost at the same usage
Llama Llama Llama 3.3 70B Least filtered Tinfoil model, but weak
Mistral / cheap vision Mistral Gemma 4 31B Cheap, reads images
Claude, Gemini, Grok as the real models n/a None Disable those slots, or show "not available in Private mode"

For low-filter chats, keep a Venice toggle (directly, or through OpenRouter's Venice endpoints). Tinfoil won't replace it.

Estimated monthly cost: one heavy user (1,200 messages × 6K tokens in, 400 out)

That's 7.2M input tokens and 0.48M output tokens a month. Prices come from the catalog feed. These are my calculations.

Setup Monthly cost
All DeepSeek V4.1 Flash $5.38 (about $3.50 if half the input hits the prompt cache)
All GLM-5.3 $15.72 (about $10.86 with 50% cache hits)
All Kimi K3 $38.40
All gpt-oss-120b (cheapest, strict) $1.37
Recommended mix: 70% DeepSeek V4.1 Flash, 20% GLM-5.3, 10% Kimi K3 ≈ $10.75 base. Range about $7–$14: about $7.31 with 50% cache hits, about $13.75 if each answer adds roughly 600 hidden reasoning tokens

Cost assumptions (Unverified): the 50% cache-hit rate and 600 reasoning tokens are illustrative guesses, not measurements. Real numbers depend on how much chat history repeats and on the reasoning effort setting. For comparison, the same 1,200 messages on Venice's Grok 4.7 would cost about $19.61 ($2.27 / $6.80 per 1M, Venice models API).

Practical tip: set a monthly spend limit (for example $20) on the Tinfoil key so the toggle can never surprise you (get an API key).

10. Sources (all accessed Oct 6, 2026)

Tinfoil, official: - tinfoil.sh - Pricing - Model catalog, with its feed at api.tinfoil.sh/api/config/models - Private Inference - Privacy policy (Sept 22, 2026) - Security & Privacy FAQ (Sept 25, 2026) - Terms - Acceptable use policy - Safety & Safeguards - Company - Customers - Technology - Blog, especially the confidential-computing overhead post - Status page and status RSS

Tinfoil docs: - Docs index (llms.txt) - Models: Chat · Vision · Audio · Embeddings · Safety · Overview - Getting started: API key · Python SDK · Direct API · Proxy CLI - Guides: Reasoning · Tool calling · Structured outputs · Prompt caching · Errors · Admin API - Security: Verification · Architecture · Enclave primer · EHBP - Changelog

Tinfoil code: - github.com/tinfoilsh - safeguard-evals · confidential-safeguards · confidential-model-router - Model configs: Kimi K3 · GLM-5.3 · DeepSeek V4.1 Flash - PyPI: tinfoil

Community and press: - Launch HN - r/LocalLLaMA thread - Y Combinator profile - Funding databases: Dealroom · StartupHub · LinkedIn - chatgate (Product Hunt summary) - Duck.ai: privacy help page · Duck.ai help

Price and directory trackers: - endpoints.run/providers/tinfoil - confidentialinference.net - llmprice

Benchmarks: - Artificial Analysis: leaderboard · Intelligence Index · GPQA · HLE · SciCode - AA model pages: Gemma 4 31B · gpt-oss-120b · Llama 3.3 70B · Grok 4.7 · DeepSeek V4.1 Flash - LMArena text - LiveBench: via BenchLeader · livebench.ai

Venice: - Models API - Privacy docs - Pricing - Terms of service

OpenRouter: - Providers list - Endpoints: GLM-5.3 · Kimi K3 · DeepSeek V4.1 Flash

Privatemode: - Pricing - Imprint - Inference API

Scaleway: - Data privacy - Pricing

Prior research: privacy-llm-providers-2026-10-06.md, Chance AI research folder, Oct 6, 2026.