# FuseLLM > FuseLLM is a free, browser-only AI workspace. Bring your own API keys for Claude, ChatGPT, Gemini, Grok, DeepSeek, Kimi, Qwen, GLM, MiniMax, Nemotron and Perplexity Sonar, then wire them into circuits where one model builds, another reviews, and the loop runs until the work is done. Circuits can generate images, video and music, push code to GitHub, send email and post to Slack. No server, no account, keys never leave your device. FuseLLM is at https://fusellm.lowkey.tools/. Built by [Shrinath Prabhu](https://shrinath.me) ([@shrinath_prabhu](https://x.com/shrinath_prabhu)), from the makers of [OwlEye Analytics](https://owleye.dev), and published on [lowkey.tools](https://lowkey.tools). ## What it does - **Bring your own keys.** Paste an OpenRouter key and every model is live. Or use a Perplexity key, or direct keys for OpenAI, Anthropic, Google, DeepSeek, xAI, Moonshot, Qwen and MiniMax. Every key has an ⓘ with where to get it and how to cap it. Keys stay in this browser and can be locked with a passphrase. - **Circuits, not prompts.** Wire models together like a Zapier zap. Claude Fable writes the code, GPT-6 Astra reviews it, and the review loops back until the reviewer approves. No approval clicks in between. - **Four kinds of wire.** Each connection can carry the Output of the last step, the original Input, the full Context of every step so far, and a shared Memory that any model can write to. Mix all four. - **Roles and skills.** Give every model a role, from senior engineer to data analyst, teacher, copywriter or chief of staff, and stack skills such as code review, SQL, translation, meeting notes or red teaming. Twenty-one roles and twenty-two skills ship built in, and every one can be edited, deleted or cloned. - **MCP tools in the browser.** Eight remote MCP servers are set up and waiting: DeepWiki, Context7, GitHub, Notion, Jira and Confluence, Linear, Asana and Hugging Face. Models call their tools mid-answer, and this is how apps that refuse browser requests are reached. Web search is one switch away. - **Every token on the meter.** Live thinking and generating labels, elapsed time and token counts on every step, like a coding agent in a terminal. Totals and cost at the end. Set a stop-loss and the run halts before it spends more. - **Fast, balanced or deep.** One switch tunes reasoning effort, output length and tone per model. Fast for quick drafts, deep think when correctness matters more than speed. - **Apps and actions, like Zapier.** End a circuit by pushing code to GitHub, emailing from Gmail, Zoho or Outlook, filling a Google Doc, Sheet or Slides deck, filing tasks in Asana, Todoist, Trello, Linear or Jira, deploying to Vercel or Netlify, commenting on Figma, saving to Drive or Dropbox, or publishing the video to YouTube. Eighteen apps, twenty-eight actions, and models can call them as tools too. - **Images, video, music and voice.** The Studio and media stages reach 50+ image models, Veo, Sora, Kling, Runway, Hailuo and Seedance video, Lyria music and a dozen voices on the same OpenRouter key, plus ElevenLabs music, sound effects and voices, and fal.ai. A vision model can critique an image and loop it back for edits. - **A film studio in a circuit.** Script, review, direction, a character sheet, a keyframe and a video clip per shot, narration and score, an audio check, then a final cut with titles and crossfades rendered in your browser, and a screening that sends notes back to the edit. - **Token calculator.** Paste any text to count its tokens exactly, see what it costs on every model, and estimate a whole chat or circuit before you run it: wires, loops, reasoning and all. One click tidies a prompt (typically 20-50% smaller, code and links untouched), and a reply cap cuts the output side, where the money actually goes. - **Sources you can check.** When Sonar, Claude web search or OpenRouter search report the pages they used, FuseLLM lists them under the answer with numbered citations, carries them into later stages, and keeps them in exports. - **Research notes into Superbrain.** Export any run or chat as a Superbrain vault: linked Markdown notes with sources, memory and images, ready to open in the Superbrain notes app on lowkey.tools. - **Installable and offline-first.** A mobile-first progressive web app. It opens with no connection, so your chats, circuits and library are always there. Mirror everything to a folder on your computer, and share circuits as files. Talking to a model needs the internet, nothing else does. ## How it works 1. **Add a key.** Open Models and paste an OpenRouter key, or a direct key from any supported provider. Enable the models you want. 2. **Pick or build a circuit.** Start from a template such as "Code, review, repeat" or "Student and professor", or add stages yourself. Each stage is a model with a role, skills and tools. 3. **Wire the stages.** Choose what flows into each stage: Output, Input, Context, Memory. Add a loop so a reviewer can send work back until it approves. 4. **Run it.** Type the brief and press Run. Watch each stage think and generate with live time and token counts, then copy or download the final result. ## Supported models - GPT-6 Astra (OpenAI): OpenAI flagship. Strongest at hard reasoning, agentic coding and careful review. - GPT-5.6 Sol (OpenAI): Balanced all-rounder for planning, writing and everyday code. - GPT-5.6 Terra (OpenAI): Grounded analysis and research synthesis at a mid-tier price. - GPT-5.6 Luna (OpenAI): Small, quick and cheap. Good for drafts, summaries and triage steps. - Claude Fable 5.1 (Anthropic): Anthropic's most capable model. Long-horizon coding and the hardest reasoning. - Claude Opus 5 (Anthropic): Deep reasoning and meticulous review. A natural professor or architect. - Claude Sonnet 5 (Anthropic): Fast, capable and affordable. Great builder, editor and technical writer. - Gemini 3.8 Flash (Google): Fast, cheap, a million tokens of context. A tireless research assistant. - Kimi K3 (Moonshot AI): Agentic coder with long context and strong tool use. - DeepSeek V4 Pro (DeepSeek): Frontier-level reasoning at a fraction of the price. A sharp critic. - Grok 4.6 (xAI): Direct, contrarian and current. Good debater and fact checker. - GLM 5.3 (Z.ai): Open-weight coder with parallel tool calls. OpenRouter only from a browser. - Qwen 3.8 Flash (Alibaba): Very cheap and quick. Ideal for summaries, drafts and routing steps. - Nemotron 3 Ultra (NVIDIA): Open reasoning model via OpenRouter. The free tier costs nothing to try. - MiniMax M3 (MiniMax): Cheap, long-context agent model. Solid second opinion on a budget. - Perplexity Sonar Pro (Perplexity): Answers from the live web with citations on every claim. The fact-finder of any research circuit. - Perplexity Sonar Deep Research (Perplexity): Runs dozens of searches and reads the sources before it writes. Slow, thorough, cited. ## Facts - Price: free. Users pay their own AI provider for tokens, through their own key. - Account: none. No sign-up, no email, no login. - Backend: none. The browser talks to AI providers directly; data stays in IndexedDB on the device. - Keys: stored only in the browser, optionally encrypted with a passphrase (AES-GCM, PBKDF2). - Offline: installable PWA that opens offline; running a model needs a connection. - Circuits: automated multi-model chains with four wire types (Input, Output, Context, Memory) and review loops that repeat until a verdict of APPROVED. - Metering: live thinking and generating status, elapsed time and token counts per step, totals and cost per run, and an optional token stop-loss. - Modes: Fast, Balanced and Deep think, which tune reasoning effort and answer length per provider. - Tools: remote MCP servers over Streamable HTTP, and web search. ## Answers - **What is FuseLLM?** FuseLLM is a free AI workspace that runs entirely in your browser. You bring your own API keys, chat with leading models, and wire several models into circuits where each one builds on, reviews or extends the work of the others until the task is done. - **Is FuseLLM free?** Yes. FuseLLM itself is free and has no paid tier. You only pay your AI provider for the tokens you use, at their normal rates, through your own key. Some OpenRouter models, such as Nemotron 3 Ultra (free), cost nothing. - **Are my API keys safe?** Your keys are stored only in this browser and are sent only to the AI provider they belong to. FuseLLM has no backend, so there is no server that could see them. You can lock them behind a passphrase, which encrypts them with AES-GCM on your device. - **What is a circuit in FuseLLM?** A circuit is an automated chain of AI models, like a Zapier zap for LLMs. Each stage is a model with a role, skills and tools. Wires pass the output, the original input, the full context or shared memory from one stage to the next, and loops let a reviewer send work back until it is approved. - **Which AI models does FuseLLM support?** GPT-6 Astra, GPT-5.6 Sol, Terra and Luna from OpenAI; Claude Fable 5.1, Opus 5 and Sonnet 5 from Anthropic; Gemini 3.8 Flash; Kimi K3; DeepSeek V4 Pro; Grok 4.6; GLM 5.3; Qwen 3.8 Flash; Nemotron 3 Ultra; MiniMax M3; and Perplexity Sonar Pro and Sonar Deep Research. All of them work through a single OpenRouter key, and a Perplexity key reaches Sonar and eleven of the others with web search built in. - **Can FuseLLM send emails or push code to GitHub?** Yes. Connect GitHub with a fine-grained token and a circuit can create a repository and commit every generated file in one push. Connect Gmail to send mail from your own address, or EmailJS to send through Zoho Mail, Outlook or any SMTP server. Eighteen apps are built in: GitHub, GitLab, Google Workspace (Gmail, Docs, Sheets, Slides, Drive, Calendar, YouTube), Slack, Discord, Telegram, Linear, Asana, Todoist, Trello, Airtable, Netlify, Vercel, Figma, Sentry, Dropbox and webhooks into Make, Zapier, n8n or Pipedream. Each uses a credential you create and can revoke. - **Can FuseLLM generate images, video and music?** Yes, with your OpenRouter key. The Studio and circuit media stages reach more than 50 image models including Gemini 3 Pro Image and GPT Image 2, video models such as Veo 3.1, Sora 2 Pro, Kling 3 and Runway Gen-4.5, Google Lyria for music, and a dozen text-to-speech voices. Images can be edited with reference images, and a vision model can review and loop them. - **Is it safe to put API keys into a browser app?** In FuseLLM the keys are your own and never leave your device except to go straight to the provider they belong to. There is no FuseLLM server and no app-owned secret shipped in the page. Keys sit in the browser’s IndexedDB, can be encrypted with a passphrase, and are protected by a strict Content Security Policy that allows only FuseLLM’s own scripts. Use scoped keys where you can: a spend limit on OpenRouter, a fine-grained GitHub token for chosen repos, Gmail limited to sending. - **How do I get an API key for FuseLLM?** The quickest is an OpenRouter key: sign in at openrouter.ai, add a few dollars of credit, open Keys and create one with a credit limit. One key reaches every model in FuseLLM plus image, video and audio models. In FuseLLM, the ⓘ next to each provider on the Models page gives the steps for that provider and the setting that caps a key if it ever leaks. - **How many tokens will my prompt or workflow use?** Open the token calculator in FuseLLM and paste the text. It counts tokens exactly with OpenAI’s o200k_base tokenizer on your device, shows words and characters, prices the text on every model as input and as output, and estimates a whole chat or circuit: each step’s input from its wires, loops by assumption, typical reply and reasoning lengths, and the total cost. - **Does FuseLLM show sources for research answers?** Yes. When a model reports the web pages it used, as Perplexity Sonar, Claude web search and OpenRouter web search do, FuseLLM lists them under the answer with titles, sites and numbered citations. Sources are passed to later circuit stages, available as the {{sources}} placeholder, and included in Markdown and Superbrain exports. - **Where does FuseLLM store chats and circuits?** In your browser’s IndexedDB, which can hold gigabytes, with persistent storage requested so the browser does not clear it when space runs low. On Chrome, Edge and other Chromium browsers you can also mirror everything to a folder on your computer as JSON, Markdown and media files, and load it back into another browser. API keys are never written to that folder. Circuits export as .fusellm.json files you can share. - **Can FuseLLM upload a video to YouTube or write to Notion and Jira?** YouTube yes: tick the YouTube permission when you connect Google, and a circuit can publish a generated video, private by default. Notion, Jira and Confluence block browser requests, so FuseLLM reaches them through their own remote MCP servers instead, which are set up in the MCP library. Anything else can be reached by sending a webhook to Make, Zapier, n8n or Pipedream. - **How do I make a prompt use fewer tokens?** The token calculator has a one-click optimiser. It runs on your device for free and only makes changes that cannot alter meaning: whitespace, invisible characters pasted from documents, curly quotes, filler phrases, repeated paragraphs, table padding and HTML comments, never touching code blocks, inline code or URLs. It typically removes a fifth to a half of a hand-written prompt. A reply cap saves more, because output is billed at three to five times the input price, and Fast mode cuts the hidden reasoning that is billed as output too. A model can also rewrite the prompt for you if you want it shorter still. - **Can I export FuseLLM research to Superbrain?** Yes. Any circuit run or chat exports as a Superbrain vault: a zip of linked Markdown notes with an index, one note per step, sources, shared memory and generated images. Open superbrain.lowkey.tools, choose Import a vault, and pick the zip. - **Can two AI models talk to each other in FuseLLM?** Yes. That is what circuits are for. For example, Claude Fable writes code, GPT-6 Astra reviews it, and the review goes back to Claude until Astra approves. Or one model plays a student researcher while another plays the professor who grades the work. - **Does FuseLLM need a server or an account?** No. There is no sign-up and no FuseLLM server. Your browser talks to the AI providers directly, and your chats, circuits and library are stored locally in IndexedDB. - **Does FuseLLM work offline?** The app installs as a PWA and opens offline, with all of your chats, circuits, roles and skills available. Running a model needs an internet connection, because the model runs at the provider. - **How do I limit how many tokens a circuit spends?** Set a stop-loss on a chat or a circuit. FuseLLM caps each request to the budget that is left, watches tokens as they stream, and stops the run when the limit is reached. It can also squeeze context to fit. Providers bill hidden reasoning, so the limit is enforced as closely as the provider allows but is not guaranteed to the token. - **Can I use MCP servers in the browser?** Yes. FuseLLM speaks MCP over Streamable HTTP, so any remote MCP server that allows browser requests works, including DeepWiki and Context7. Attach a server to a chat or a circuit stage and the model can call its tools. - **Who makes FuseLLM?** FuseLLM is built by Shrinath Prabhu (shrinath.me) and published on lowkey.tools by the makers of OwlEye Analytics (owleye.dev), a privacy-first, cookie-free web analytics product. ## Links - [FuseLLM](https://fusellm.lowkey.tools/): the app - [Source on GitHub](https://github.com/shrinathprabhu/fusellm): source code and issues - [llms-full.txt](https://fusellm.lowkey.tools/llms-full.txt): every detail, including built-in roles, skills and circuit templates ## More from the same shelf FuseLLM is one of twelve browser-only tools on [lowkey.tools](https://lowkey.tools), built by [Shrinath Prabhu](https://shrinath.me) ([@shrinath_prabhu](https://x.com/shrinath_prabhu)) from the makers of [OwlEye Analytics](https://owleye.dev). - [SuperBrain](https://superbrain.lowkey.tools/): A private, local-first workspace for notes, links and ideas: Notion and Obsidian with none of the ceremony. - [SuperSplit](https://supersplit.lowkey.tools/): Split group expenses, track who paid and settle up fairly. No accounts, no ads, no limits. - [SuperFocus](https://superfocus.lowkey.tools/): An offline focus space: Pomodoro sessions, tasks, goals and quiet music. - [StreakFreak](https://streakfreak.lowkey.tools/): A private, offline habit tracker with templates, flexible goals and streaks worth keeping. - [Credo](https://credo.lowkey.tools/): Share passwords, secrets and small files as encrypted links that expire on their own. - [Converteasy](https://converteasy.lowkey.tools/): Units, currencies, crypto denominations, dates and calculations in one intelligent box. - [Favigen](https://favigen.lowkey.tools/): One SVG or PNG in, every favicon, app icon, manifest and HTML tag out. - [Billgen](https://billgen.lowkey.tools/): Invoices, receipts, memos and bills made locally, then exported, printed or shared. - [MathMagician](https://mathmagician.lowkey.tools/): Race the clock through as many arithmetic problems as you can, then share your best score. - [Chesscape](https://chesscape.lowkey.tools/): Escape today’s near-checkmate position with the one saving move, in thirty seconds. - [SpotFast](https://spotfast.lowkey.tools/): Memorise a grid in seconds, then find the hidden targets before your three lives run out. - [OwlEye Analytics](https://owleye.dev): Privacy-first, cookie-free web analytics. - [Shrinath Prabhu](https://shrinath.me): the developer - [lowkey.tools](https://lowkey.tools): more small, free tools from the same makers