FuseLLM

Free · Bring your own keys · Runs in your browser

Wire AI models into circuits
that finish the job.

A free, browser-only BYOK AI workspace. Wire the best models into self-running circuits that research, build, review, make images and video, and ship to GitHub, Gmail or Slack.

One OpenRouter key unlocks all 17 models and the Studio, including a free one. Keys never leave this device.

What it does

  • Bring your own keys

    Paste an OpenRouter key and every model is live. Or use a Perplexity key, or direct keys for OpenAI, Anthropic, Google, DeepSeek, xAI, Moonshot, Qwen and MiniMax. Every key has an ⓘ with where to get it and how to cap it. Keys stay in this browser and can be locked with a passphrase.

  • Circuits, not prompts

    Wire models together like a Zapier zap. Claude Fable writes the code, GPT-6 Astra reviews it, and the review loops back until the reviewer approves. No approval clicks in between.

  • Four kinds of wire

    Each connection can carry the Output of the last step, the original Input, the full Context of every step so far, and a shared Memory that any model can write to. Mix all four.

  • Roles and skills

    Give every model a role, from senior engineer to data analyst, teacher, copywriter or chief of staff, and stack skills such as code review, SQL, translation, meeting notes or red teaming. Twenty-one roles and twenty-two skills ship built in, and every one can be edited, deleted or cloned.

  • MCP tools in the browser

    Eight remote MCP servers are set up and waiting: DeepWiki, Context7, GitHub, Notion, Jira and Confluence, Linear, Asana and Hugging Face. Models call their tools mid-answer, and this is how apps that refuse browser requests are reached. Web search is one switch away.

  • Every token on the meter

    Live thinking and generating labels, elapsed time and token counts on every step, like a coding agent in a terminal. Totals and cost at the end. Set a stop-loss and the run halts before it spends more.

  • Fast, balanced or deep

    One switch tunes reasoning effort, output length and tone per model. Fast for quick drafts, deep think when correctness matters more than speed.

  • Apps and actions, like Zapier

    End a circuit by pushing code to GitHub, emailing from Gmail, Zoho or Outlook, filling a Google Doc, Sheet or Slides deck, filing tasks in Asana, Todoist, Trello, Linear or Jira, deploying to Vercel or Netlify, commenting on Figma, saving to Drive or Dropbox, or publishing the video to YouTube. Eighteen apps, twenty-eight actions, and models can call them as tools too.

  • Images, video, music and voice

    The Studio and media stages reach 50+ image models, Veo, Sora, Kling, Runway, Hailuo and Seedance video, Lyria music and a dozen voices on the same OpenRouter key, plus ElevenLabs music, sound effects and voices, and fal.ai. A vision model can critique an image and loop it back for edits.

  • A film studio in a circuit

    Script, review, direction, a character sheet, a keyframe and a video clip per shot, narration and score, an audio check, then a final cut with titles and crossfades rendered in your browser, and a screening that sends notes back to the edit.

  • Token calculator

    Paste any text to count its tokens exactly, see what it costs on every model, and estimate a whole chat or circuit before you run it: wires, loops, reasoning and all. One click tidies a prompt (typically 20-50% smaller, code and links untouched), and a reply cap cuts the output side, where the money actually goes.

  • Sources you can check

    When Sonar, Claude web search or OpenRouter search report the pages they used, FuseLLM lists them under the answer with numbered citations, carries them into later stages, and keeps them in exports.

  • Research notes into Superbrain

    Export any run or chat as a Superbrain vault: linked Markdown notes with sources, memory and images, ready to open in the Superbrain notes app on lowkey.tools.

  • Installable and offline-first

    A mobile-first progressive web app. It opens with no connection, so your chats, circuits and library are always there. Mirror everything to a folder on your computer, and share circuits as files. Talking to a model needs the internet, nothing else does.

How it works

  1. Add a key

    Open Models and paste an OpenRouter key, or a direct key from any supported provider. Enable the models you want.

  2. Pick or build a circuit

    Start from a template such as "Code, review, repeat" or "Student and professor", or add stages yourself. Each stage is a model with a role, skills and tools.

  3. Wire the stages

    Choose what flows into each stage: Output, Input, Context, Memory. Add a loop so a reviewer can send work back until it approves.

  4. Run it

    Type the brief and press Run. Watch each stage think and generate with live time and token counts, then copy or download the final result.

Circuits people run

  • Code, review, repeat. Claude Fable writes it, GPT-6 Astra reviews it, and the loop runs until the review passes.
  • Student and professor. Gemini drafts research with web search, Claude Opus grades it and sends it back with notes.
  • Plan to product. Spec, architecture, implementation, review, tests and docs, one model per job, one brief to start it.
  • Deep research. Break a question down, research each part, fact-check the claims, then write the report.
  • Debate and judge. Grok argues for, DeepSeek argues against, and Opus weighs the case.
  • Build and ship. Plan, build and review until approved, then commit the code to a new GitHub repo in one push.
  • Art director loop. Claude writes the prompt, Gemini paints it, a vision model critiques it and sends edits back.
  • Podcast in a box. Sonar researches, Sonnet scripts, a voice model narrates and Lyria scores the intro.
  • Inbox to next actions. Paste a morning of email and messages; get today’s plan and the tasks filed in Todoist.
  • Error triage. Pull unresolved Sentry errors, group them by root cause, write the fix, open the GitHub issue.
  • Meeting to memo. A transcript becomes decisions, owners and dates, checked against the transcript and saved as a Google Doc.
  • Landing page, live. Copy, a built page, an accessibility review, then a deploy that returns the URL.
  • Movie studio. Script to screening in twelve stages: characters, keyframes, Veo clips per shot, narration, score and a final cut, with review loops at every step.

Supported models

  • GPT-6 Astra · OpenAI
  • GPT-5.6 Sol · OpenAI
  • GPT-5.6 Terra · OpenAI
  • GPT-5.6 Luna · OpenAI
  • Claude Fable 5.1 · Anthropic
  • Claude Opus 5 · Anthropic
  • Claude Sonnet 5 · Anthropic
  • Gemini 3.8 Flash · Google
  • Kimi K3 · Moonshot AI
  • DeepSeek V4 Pro · DeepSeek
  • Grok 4.6 · xAI
  • GLM 5.3 · Z.ai
  • Qwen 3.8 Flash · Alibaba
  • Nemotron 3 Ultra · NVIDIA
  • MiniMax M3 · MiniMax
  • Perplexity Sonar Pro · Perplexity
  • Perplexity Sonar Deep Research · Perplexity

Questions

What is FuseLLM?

FuseLLM is a free AI workspace that runs entirely in your browser. You bring your own API keys, chat with leading models, and wire several models into circuits where each one builds on, reviews or extends the work of the others until the task is done.

Is FuseLLM free?

Yes. FuseLLM itself is free and has no paid tier. You only pay your AI provider for the tokens you use, at their normal rates, through your own key. Some OpenRouter models, such as Nemotron 3 Ultra (free), cost nothing.

Are my API keys safe?

Your keys are stored only in this browser and are sent only to the AI provider they belong to. FuseLLM has no backend, so there is no server that could see them. You can lock them behind a passphrase, which encrypts them with AES-GCM on your device.

What is a circuit in FuseLLM?

A circuit is an automated chain of AI models, like a Zapier zap for LLMs. Each stage is a model with a role, skills and tools. Wires pass the output, the original input, the full context or shared memory from one stage to the next, and loops let a reviewer send work back until it is approved.

Which AI models does FuseLLM support?

GPT-6 Astra, GPT-5.6 Sol, Terra and Luna from OpenAI; Claude Fable 5.1, Opus 5 and Sonnet 5 from Anthropic; Gemini 3.8 Flash; Kimi K3; DeepSeek V4 Pro; Grok 4.6; GLM 5.3; Qwen 3.8 Flash; Nemotron 3 Ultra; MiniMax M3; and Perplexity Sonar Pro and Sonar Deep Research. All of them work through a single OpenRouter key, and a Perplexity key reaches Sonar and eleven of the others with web search built in.

Can FuseLLM send emails or push code to GitHub?

Yes. Connect GitHub with a fine-grained token and a circuit can create a repository and commit every generated file in one push. Connect Gmail to send mail from your own address, or EmailJS to send through Zoho Mail, Outlook or any SMTP server. Eighteen apps are built in: GitHub, GitLab, Google Workspace (Gmail, Docs, Sheets, Slides, Drive, Calendar, YouTube), Slack, Discord, Telegram, Linear, Asana, Todoist, Trello, Airtable, Netlify, Vercel, Figma, Sentry, Dropbox and webhooks into Make, Zapier, n8n or Pipedream. Each uses a credential you create and can revoke.

Can FuseLLM generate images, video and music?

Yes, with your OpenRouter key. The Studio and circuit media stages reach more than 50 image models including Gemini 3 Pro Image and GPT Image 2, video models such as Veo 3.1, Sora 2 Pro, Kling 3 and Runway Gen-4.5, Google Lyria for music, and a dozen text-to-speech voices. Images can be edited with reference images, and a vision model can review and loop them.

Is it safe to put API keys into a browser app?

In FuseLLM the keys are your own and never leave your device except to go straight to the provider they belong to. There is no FuseLLM server and no app-owned secret shipped in the page. Keys sit in the browser’s IndexedDB, can be encrypted with a passphrase, and are protected by a strict Content Security Policy that allows only FuseLLM’s own scripts. Use scoped keys where you can: a spend limit on OpenRouter, a fine-grained GitHub token for chosen repos, Gmail limited to sending.

How do I get an API key for FuseLLM?

The quickest is an OpenRouter key: sign in at openrouter.ai, add a few dollars of credit, open Keys and create one with a credit limit. One key reaches every model in FuseLLM plus image, video and audio models. In FuseLLM, the ⓘ next to each provider on the Models page gives the steps for that provider and the setting that caps a key if it ever leaks.

How many tokens will my prompt or workflow use?

Open the token calculator in FuseLLM and paste the text. It counts tokens exactly with OpenAI’s o200k_base tokenizer on your device, shows words and characters, prices the text on every model as input and as output, and estimates a whole chat or circuit: each step’s input from its wires, loops by assumption, typical reply and reasoning lengths, and the total cost.

Does FuseLLM show sources for research answers?

Yes. When a model reports the web pages it used, as Perplexity Sonar, Claude web search and OpenRouter web search do, FuseLLM lists them under the answer with titles, sites and numbered citations. Sources are passed to later circuit stages, available as the {{sources}} placeholder, and included in Markdown and Superbrain exports.

Where does FuseLLM store chats and circuits?

In your browser’s IndexedDB, which can hold gigabytes, with persistent storage requested so the browser does not clear it when space runs low. On Chrome, Edge and other Chromium browsers you can also mirror everything to a folder on your computer as JSON, Markdown and media files, and load it back into another browser. API keys are never written to that folder. Circuits export as .fusellm.json files you can share.

Can FuseLLM upload a video to YouTube or write to Notion and Jira?

YouTube yes: tick the YouTube permission when you connect Google, and a circuit can publish a generated video, private by default. Notion, Jira and Confluence block browser requests, so FuseLLM reaches them through their own remote MCP servers instead, which are set up in the MCP library. Anything else can be reached by sending a webhook to Make, Zapier, n8n or Pipedream.

How do I make a prompt use fewer tokens?

The token calculator has a one-click optimiser. It runs on your device for free and only makes changes that cannot alter meaning: whitespace, invisible characters pasted from documents, curly quotes, filler phrases, repeated paragraphs, table padding and HTML comments, never touching code blocks, inline code or URLs. It typically removes a fifth to a half of a hand-written prompt. A reply cap saves more, because output is billed at three to five times the input price, and Fast mode cuts the hidden reasoning that is billed as output too. A model can also rewrite the prompt for you if you want it shorter still.

Can I export FuseLLM research to Superbrain?

Yes. Any circuit run or chat exports as a Superbrain vault: a zip of linked Markdown notes with an index, one note per step, sources, shared memory and generated images. Open superbrain.lowkey.tools, choose Import a vault, and pick the zip.

Can two AI models talk to each other in FuseLLM?

Yes. That is what circuits are for. For example, Claude Fable writes code, GPT-6 Astra reviews it, and the review goes back to Claude until Astra approves. Or one model plays a student researcher while another plays the professor who grades the work.

Does FuseLLM need a server or an account?

No. There is no sign-up and no FuseLLM server. Your browser talks to the AI providers directly, and your chats, circuits and library are stored locally in IndexedDB.

Does FuseLLM work offline?

The app installs as a PWA and opens offline, with all of your chats, circuits, roles and skills available. Running a model needs an internet connection, because the model runs at the provider.

How do I limit how many tokens a circuit spends?

Set a stop-loss on a chat or a circuit. FuseLLM caps each request to the budget that is left, watches tokens as they stream, and stops the run when the limit is reached. It can also squeeze context to fit. Providers bill hidden reasoning, so the limit is enforced as closely as the provider allows but is not guaranteed to the token.

Can I use MCP servers in the browser?

Yes. FuseLLM speaks MCP over Streamable HTTP, so any remote MCP server that allows browser requests works, including DeepWiki and Context7. Attach a server to a chat or a circuit stage and the model can call its tools.

Who makes FuseLLM?

FuseLLM is built by Shrinath Prabhu (shrinath.me) and published on lowkey.tools by the makers of OwlEye Analytics (owleye.dev), a privacy-first, cookie-free web analytics product.

The rest of the shelf

FuseLLM is one of twelve tools on lowkey.tools. Same idea every time: open the page, do the thing, no account, nothing kept on a server.