Show HN: Mysti – Claude, Codex, and Gemini debate your code, then synthesize

Posted by bahaAbunojaim 12/23/2025

Show HN: Mysti – Claude, Codex, and Gemini debate your code, then synthesize(github.com)

Hey HN! I'm Baha, creator of Mysti.

The problem: I pay for Claude Pro, ChatGPT Plus, and Gemini but only one could help at a time. On tricky architecture decisions, I wanted a second opinion.

The solution: Mysti lets you pick any two AI agents (Claude Code, Codex, Gemini) to collaborate. They each analyze your request, debate approaches, then synthesize the best solution.

Your prompt → Agent 1 analyzes → Agent 2 analyzes → Discussion → Synthesized solution

Why this matters: each model has different training and blind spots. Two perspectives catch edge cases one would miss. It's like pair programming with two senior devs who actually discuss before answering.

What you get: * Use your existing subscriptions (no new accounts, just your CLI tools) * 16 personas (Architect, Debugger, Security Expert, etc) * Full permission control from read-only to autonomous * Unified context when switching agents

Tech: TypeScript, VS Code Extension API, shells out to claude-code/codex-cli/gemini-cli

License: BSL 1.1, free for personal and educational use, converts to MIT in 2030 (would love input on this, does it make sense to just go MIT?)

GitHub: https://github.com/DeepMyst/Mysti

Would love feedback on the brainstorm mode. Is multi-agent collaboration actually useful or am I just solving my own niche problem?

216 points | 178 comments

d4rkp4ttern 12/27/2025|

A workflow I find useful is to have multiple CLI agents running in different Tmux panes and have one consult/delegate to another using my Tmux-CLI [1] tool + skill. Advantage of this is that the agents’ work is fully visible and I can intervene as needed.

[1] https://github.com/pchalasani/claude-code-tools?tab=readme-o...

vidarh 12/27/2025||

Have you considered using their command line options instead? At least Codex and Claude both support feeding in new prompts in an ongoing conversation via the command line, and can return text or stream JSON back.

d4rkp4ttern 12/27/2025||

You mean so-called headless or non-interactive mode? Yes I’ve considered that but the advantage communication via Tmux panes is that all agent work is fully visible and you can intervene as needed.

My repo has other tools that leverage such headless agents; for example there’s a resume [1] functionality that provides alternatives to compaction (which is not great since it always loses valuable context details): The “smart-trim” feature uses a headless agent to find irrelevant long messages for truncation, and the “rollover” feature creates a new session and injects session lineage links, with a customizable extraction of context for the task to be continued.

[1] https://github.com/pchalasani/claude-code-tools?tab=readme-o...

petesergeant 12/27/2025|||

I've had good success with a similar workflow, most recently using it to help me build out a captive-wifi debugger[0]. In short, it worked _pretty_ well, but it was quite time intensive. That said, I think removing the human from the loop would have been insanity on this: lots of situations where there were some very poor ideas suggested that the other LLMs went along with, and others where one LLM was the sole voice of reason against the other two.

I think my only real take-away from all of it was that Claude is probably the best at prototyping code, where Codex make a very strong (but pedantic) code-reviewer. Gemini was all over the place, sometimes inspired, sometimes idiotic.

0: https://github.com/pjlsergeant/captive-wifi-tool/tree/main

bahaAbunojaim 12/27/2025||

This is exactly why I built Mysti because I used that flow very often and it worked well, I also added personas and skills so that it is easy to customize the agents behavior and if you have any ideas to make the behavior better then please don’t hesitate to share! Happy to jump on a call and discuss it as well

bikeshaving 12/27/2025|||

I have a similar workflow except I haven’t put time into the tooling - Claude is adept at TMUX and it can almost even prompt and respond to ChatGPT except it always forgets to press Enter when it sends keys. Have your agents been able to communicate with each other with tmux send-keys?

theturtletalks 12/27/2025|||

I had the same issue. Subagents are nice but the LLM calling them can’t have a back and forth conversation. I tried tmux-cli and even other options like AgentAPI[0] but the same issue persists, the agent can’t have a back and forth with the tmux pane.

To people asking why would you want Claude to call Codex or Gemini, it’s because of orchestration. We have an architect skill we feed the first agent. That agent can call subagents or even use tmux and feed in the builder skill. The architect is harnessed to a CRUD application just keeping track of what features were built already so the builder is focused on building only.

0. https://github.com/coder/agentapi

d4rkp4ttern 12/27/2025||||

Yes this and other edge cases is why I made the Tmux-CLI wrapper. Yes they use send-keys with suitable delays etc

zingar 12/27/2025|||

What are you asking/expecting Claude to do with tmux?

bikeshaving 12/27/2025||

I find that asking Claude to develop and Codex to review the uncommitted changes will typically result in high-value code, and eliminate all of Claude’s propensity to perpetually lie and cheat. Sometimes I also ideate with Claude and then ask Claude to get ChatGPT’s opinion on the matter. I started by copy-pasting responses but I found tmux to be a nice way to get rid of the middleman.

joshstrange 12/28/2025||

What does tmux add here? Or how does it allow you to do that? I’m sorry I’m just missing it I’m sure. I don’t use tmux a lot so I don’t know all its potential.

aoeusnth1 12/29/2025||

It lets Claude directly type into Codex as if it were the user, or vice versa

zingar 12/29/2025||

And you’re finding that it can’t do that without tmux?

bahaAbunojaim 12/27/2025|||

I will look it up indeed

wild_egg 12/28/2025|||

What does Tmux-CLI add on top of regular tmux?

Everything in the "What Claude Code Can Do With Tmux-CLI" section is already easily possible out of the box with vanilla tmux

d4rkp4ttern 12/28/2025||

You're right that vanilla tmux can do all of this, if a human were to use it. tmux-cli exists because LLMs frequently make mistakes with raw tmux: forgetting the Enter key, not adding delays between text and Enter (causing race conditions with fast CLI apps), or incorrect escaping.

It bakes in defaults that address these: Enter is sent automatically with a 1-second delay (configurable), pane targeting accepts simple numbers instead of session:window.pane, and there's built-in wait_idle to detect when a CLI is ready for input. Basically a wrapper that eliminates the common failure modes I kept hitting when having Claude Code interact with other terminal sessions.

throwaway12345t 12/27/2025|||

This is cool, if Codex or Gemini CLI is supported it would be good to have a section in the readme indicating shortcomings etc (may have missed)

tikimcfee 12/27/2025|||

The idea works well with or without direct integration. You can have a cli agent read arbitrary state of any tmux session and have it drive work through it. I use it for everything from dev work to system debugging. It turns out a portable and callable binary with simple parameters is still easier to use for agents than protocols and skills: https://github.com/tikimcfee/gomuxai

d4rkp4ttern 12/27/2025||||

There’s no special support needed; it’s just a bash command that any CLI agent can use. For agents that have skills, the corresponding skill helps leverage more easily. I’ll add that to the README

bahaAbunojaim 12/27/2025|||

Claude code, Gemini and codex are all supported but need more testing so I would really value the feedback, bug reports and contributions as well :D

Contributions will be highly appreciated and credited

sharifabdel 12/27/2025||

What prompted you to build this?

d4rkp4ttern 12/27/2025|||

I have both Codex and Claude subs so I wanted one to be able to consult the other. Also it’s useful when you have a cli script that an agent is iterating on, so it can test it. Another use case is for a CLI agent to run a debugger like PDB in another pane, though I haven’t used it much.

bahaAbunojaim 12/27/2025|||

I used to get stuck sometimes with Claude and needing a different agent to take a look and the switch back and forth between those agents is a headache and also you won’t be able to port all the context so thought this might help solve real blockers for many devs on larger projects

csar 12/27/2025||

Getting feedback on a plan or implementation is valuable because you get a fresh set of eyes. Using multiple models may help though it always feels a bit silly to me (if nothing else you’re increasing non-determinism because you know have to understand 2 LLM’s quirks).

But the “playing house” approach of experts is somewhere between pointless and actively harmful. It was all the rage in June and I thought people abandoned that later in the summer.

If you want the model to eg review code instead of fixing things, or document code without suggesting improvements (for writing docs), that’s useful. But there’s. I need for all these personas.

bahaAbunojaim 12/27/2025|

The way it works is that each agent think independently, discuss the solution and each agent opinion then one will synthesize a solution.

csar 12/27/2025||

I understand. My point is that the personas are generally not a good idea and that there are much simpler and more predictable ways of getting better results.

jacob019 12/28/2025||

I get where you're coming from, especially since role playing was so vital in early models in a way that is no longer necessary, or even harmful; however, when designing a complex system of interactions, there's really no way around it. And as humans we do this constantly, putting on a different hat for different jobs. When I'm wearing my developer hat, I have to reason about the role of each component in a system, and when I use an agent to serve in that role, by curating it's context and designating rules for how I want it to behave, I'm assigning it a persona. What's more, I may prime the context user and assistant messages, as examples of how I want it to respond. That context becomes the agent's personality--it's persona.

bahaAbunojaim 12/28/2025||

Spot on

tombert 12/28/2025||

I so want to like these vibe coding agents, and sometimes I do, but it really does kind of suck the joy out of things.

What I was hoping would be that I could effectively farm out work to my metaphorical AI intern while I get to focus on fun and/or interesting work. Sometimes that is what happens and it makes me very happy when it does. A lot of the time, however, it generates code that is wrong, or incomplete (while claiming it is complete), and so I end up having to babysit the code, either by further prompting or just editing the code.

And then it makes a lot of software engineering become "type prompt, sit and wait a minute, look at the code, repeat", which means I'm decidedly not focusing the fun part of the project and instead I'm just larping as a manager who backseat codes.

A friend of mine said that he likes to do this backwards: he writes a lot of the code himself and then he uses Claude Code to debug and automate writing tedious stuff like unit tests, and I think that might make it a little less mind numbing.

Also, very tangential, and maybe my prompting game isn't completely on point here, but Codex seems decidedly bad at concurrent code [1]. I was working on some lock-free data store stuff, and Codex really wanted to add a bunch of lock files that were wholly unnecessary. Oh, and it kept trying to add mutexes into Rust, no matter how many times I tell it I don't want locks and it should use one-shot channels instead. To be fair, when I went and fixed the functions myself in a few spots and then told it to use that as an example, it did get a little better.

[1] I think this particular case is because it's trained on example code from Github and most code involving concurrency uses locks (incorrectly or at least sub-optimally). I guess this particular problem may be more of the fault of American universities teaching concurrent programming incorrectly at the undergrad level.

bahaAbunojaim 12/28/2025|

I find it useful to let one agent come up with a plan after a review and another agent implementing the plan. For example, Gemini reviewing the code, codex writing a plan and then Claude code implementing it

mycall 12/28/2025||

What about the reverse, after Claude code implements it, let Gemini/Codex do a code review for bugs and architecture revisions? I found it is important to prompt to only make absolutely minimal changes to the working code, or unwanted code clobbering will happen.

bahaAbunojaim 12/29/2025||

That works great too. Will be adding the ability to tag another agent in a near release

cheema33 12/27/2025||

I created a simple skill in Claude Code CLI that collaborates with Codex CLI. It is just a prompt saved in the skill format. It uses subagents as well.

Honest question. How is Mysti better than a simple Claude skill that does the same work?

achille 12/27/2025||

Could you share your skill and workflow? does claude launch codex in a tmux session?

bahaAbunojaim 12/27/2025|||

The skill would allow Claude Code CLI to call Codex CLI but then Claude Code CLI will need to pass context to Codex which would require writing the context "which causes latency" and this process of writing the context will provide limited context to Codex and also eat up from the main context window. Mysti shares the context which is very different from passing context as a parameter.

Johnny_Bonk 12/28/2025||

could you share the skill please, id lke to try it, maybe enhance it

mlrtime 12/27/2025||

Why make it a vscode extension if the point of these 3 tools is a cli interface? Meaning most of the people I know use these tools without VSCode. Is VSC required?

KronisLV 12/27/2025||

> Meaning most of the people I know use these tools without VSCode.

I guess it depends?

You can usually count on Claude Code or Codex or Gemini CLI to support the model features the best, but sometimes having a consistent UI across all of them is also nice - be it another CLI tool like OpenCode (that was a bit buggy for me when it came to copying text), or maybe Cline/RooCode/KiloCode inside of VSC, so you don't also have to install a custom editor like Cursor but can use your pre-existing VSC setup.

Okay, that was a bit of a run on sentence, but it's nice to be able to work on some context and then to switch between different models inline: "Hey Sonnet, please look at the work of the previous model up until this point and validate its findings about the cause of this bug."

I'd also love it if I could hook up some of those models (especially what Cerebras Code offers) with autocomplete so I wouldn't need Copilot either, but most of the plugins that try to do that are pretty buggy or broken (e.g. Continue.dev). KiloCode also added autocomplete, but it doesn't seem to work with BYOK.

bahaAbunojaim 12/27/2025||

Very true, I like the fact that I can now use them with a consistent UI, shared context and ability to brainstorm

Will definitely try to add those features in a future release as well

bahaAbunojaim 12/27/2025|||

That’s a great idea! I can make it a CLI too

davidmurdoch 12/27/2025||

Huh. I know hundreds that use LLMs in a VSCode based IDE, and 3 that use the CLI.

datameta 12/27/2025||

I was a proponent initially of CLI when Claude integration with VSCode required a WSL instance, but now that it is integrated directly into VSCode I feel one grouping of tooling hiccups is now ruled out in my workflow. The only (major) nitpick I have is that it wont let you finish typing and cuts you off when asking whether/how to proceed.

dwa3592 12/27/2025||

>Claude Code (Anthropic), Codex (OpenAI), and Gemini (Google) have different training, different strengths, and different blind spots.

Do they?

There was a paper about HiveMind in LLMs. They all tend to produce similar outputs when they are asked open ended questions.

monkeydust 12/27/2025||

[2510.22954] Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) https://share.google/1GHdUvhz2uhF4PVFU

bahaAbunojaim 12/27/2025|||

I usually switch agents when one agent get stuck and I faced several situations where one agent solved a problem that the other agent was stuck on

blks 12/28/2025||

It’s perceived, and perception varies across developers and time. These tools are not guaranteed to deliver anything.

spaceman_2020 12/27/2025||

I’ve never seen a profession change so fast as coding right now

blks 12/28/2025||

Don’t worry, it’s not. Just people doing busy work and spending time struggling their “tools” to make something useful.

spaceman_2020 12/28/2025||

Coding isn’t the first profession to be disrupted by automation and it won’t be the last

mycall 12/28/2025||

Wait until automation is itself automated.

CamperBob2 12/27/2025|||

Have to keep in mind that what is happening now is basically what was promised decades ago. Never mind 4GL, 5GL, expert systems, and other efforts that went nowhere... even COBOL was created with the intention of making programming look more like natural language.

Often, revolutions take longer to happen than we think they will, and then they happen faster than we think they will. And when the tipping point is finally reached, we find more people pushing back than we thought there would be.

bandrami 12/29/2025|||

OK but is it leading to either better or more plentiful software? That's the step that people keep seeming to miss here.

bahaAbunojaim 12/27/2025|||

I believe high level languages will be replaced by natural human language, the same way as low level languages replaced by high level languages. It is the natural evolution of development.

On the other hand agentic teams will take over solo agents.

Terretta 12/27/2025|||

> I believe high level languages will be replaced by natural human languageI believe high level languages will be replaced by natural human language

Ask any human client buying dev work from a web agency how "natural language" spec works out for them.

It's not clear to me at all that "natural language" alone is ideal -- unless you also have near real time iteration of builds. If you do, then the build is a concrete manifestation of the spec, and the natural language can say "but that's not what I meant".

This allows messy natural language to vector towards the solution, "I'll know it when I see it" style.

So, natural language shaping iteratively convergent builds.

blks 12/28/2025|||

Is that your actual believe? That non-formal, natural human language explaining the task will replace formal programming?

felipeerias 12/28/2025|||

Things seem to be heading in the direction of using formal languages to define deterministic behaviour and natural languages to express matters of human taste.

spaceman_2020 12/28/2025|||

If you zoom out, it does seem like the most natural thing - why should humans with finite memory and context better than an all knowing machine?

justatdotin 12/28/2025|||

I think it is actually going to happen verrrrry slowly. but it will happen. Many many of my colleagues are understandably resisting. it will take a long time to balance out.

qudat 12/28/2025|||

Meh. What I’m doing with coding agents is what I’ve been doing for years: TDD except I use prose to describe what I want instead of writing every line of code and then spend more time in review/qa

johnisgood 12/28/2025||

> except I use prose to describe what I want instead of writing every line of code

Exactly.

bandrami 12/28/2025||

And yet the output isn't noticeably different from 5 years ago.

tiku 12/27/2025||

Anyone knows of something similar but for terminal?

Update:

I've already found a solution based on a comment, and modified it a bit.

Inside claude code i've made a new agent that uses the MCP gemini through https://github.com/raine/consult-llm-mcp. this seems to work!

Claude code:

Now let me launch the Gemini MCP specialist to build the backend monitoring server:

gemini-mcp-specialist(Build monitoring backend server) ⎿ Running PreToolUse hook…

pella 12/27/2025||

https://github.com/just-every/code "Every Code - push frontier AI to it limits. A fork of the Codex CLI with validation, automation, browser integration, multi-agents, theming, and much more. Orchestrate agents from OpenAI, Claude, Gemini or any provider." Apache 2.0 ; Community fork;

ggggffggggg 12/27/2025|||

> Note: If another tool already provides a code command (e.g. VS Code), our CLI is also installed as coder. Use coder to avoid conflicts.

“If”, oh, idk, just the tool 90% of potential users will have installed.

bahaAbunojaim 12/27/2025||||

When you say orchestrate agents then what it would do? Would it allow the same context across agents and can I make agents brainstorm?

pella 12/27/2025||

  # Plan code changes (Claude, Gemini and GPT-5 consensus)
  # All agents review task and create a consolidated plan
  /plan "Stop the AI from ordering pizza at 3AM"

  # Solve complex problems (Claude, Gemini and GPT-5 race)
  # Fastest preferred (see https://arxiv.org/abs/2505.17813)
  /solve "Why does deleting one user drop the whole database?"

  # Write code! (Claude, Gemini and GPT-5 consensus)
  # Creates multiple worktrees then implements the optimal solution
  /code "Show dark mode when I feel cranky"

  # Hand off a multi-step task; Auto Drive will coordinate agents and approvals
  /auto "Refactor the auth flow and add device login"

westurner 12/28/2025|||

just-every/code: https://github.com/just-every/code ... https://news.ycombinator.com/item?id=44959671

rane 12/27/2025|||

My similar workflow within Claude Code when it gets stuck is to have it consult Gemini. Works either through Gemini CLI or the API. Surprisingly powerful pattern because I've just found that Gemini is still ahead of Opus in architectural reasoning and figuring out difficult bugs. https://github.com/raine/consult-llm-mcp

bahaAbunojaim 12/27/2025|||

This is one of the reasons I actually built it but wanted to make it more generalized to work with any agent and on the same context without switching

tiku 12/27/2025|||

I like this solution that you can ask Gemini

bahaAbunojaim 12/27/2025||

Any other ideas that you think would make it more powerful?

tiku 12/27/2025||

Perhaps that you can tell it to "use gemini for task x, claude for task y" as sub-agents.

bahaAbunojaim 12/27/2025||

How about adding the ability to tag an agent. for example:

@gemini could you review the code and then provide a summary to @claude?

@claude can you write the classes based on an architectural review by @codex

What do you think? Does that make sense ?

tikimcfee 12/27/2025|||

Here's a portable binary you drop in a directory to allow agentic cli to cross communicate with other agents, store and read state, or act as the driver of arbitrary tmux sessions in parallel: https://github.com/tikimcfee/gomuxai

bahaAbunojaim 12/27/2025||

This is very interesting, maybe I can also integrate it into Mysti

tikimcfee 12/27/2025||

Happy to help build said integration with ya, feel free to post an issue, fork, or send me a dm. The tool itself exposes the internal DB as well so others with interest can access logs, context, etc.

esafak 12/27/2025|||

http://opencode.ai/

bahaAbunojaim 12/27/2025||

Interesting indeed but would it behave the same as Claude code or will it have its own behavior, I think the system prompt is one of the key things that differentiate every agent

esafak 12/27/2025||

I do not understand your question. Even in Claude code you have access to multiple models. You can have one critique the other.

bahaAbunojaim 12/27/2025|||

I can make it for the terminal if that would be helpful, what do you think?

vulture916 12/27/2025||

Pal MCP (formerly Zen) is pretty awesome.

https://github.com/BeehiveInnovations/pal-mcp-server

bahaAbunojaim 12/27/2025||

Will give it a look indeed, I think one of the challenges with the MCP approach is that the context need to be passed and that would add to the overhead of the main agent. Is that right?

vulture916 12/27/2025||

The CLINK command will spawn separate CLI.

Don’t quote me, but I think the other methods rely on passing general detail/commands and file paths to Gemini to avoid the context overhead you’re thinking about.

danpalmer 12/28/2025||

> Together they debate, challenge each other, and synthesize the best solution

Do they? How much better are multiple agents on your evals, and what sort of evals are you running? I've also research that suggests that more agents degrades the output after a point.

bahaAbunojaim 12/28/2025|

Haven’t done Evals yet but measured on few real world situations where projects got stuck and the brainstorm mode solved it. Definitely running evals is something worth doing and contributions are welcomed

I think what really degrades the output is the context length vs context window limits, check out NoLima

danpalmer 12/29/2025||

https://www.arxiv.org/abs/2512.08296

> coordination yields diminishing or negative returns once single-agent baselines exceed ~45%

This is going to be the big thing to overcome, and without actually measuring it all we're doing is AI astrology.

bahaAbunojaim 12/29/2025||

This is why context optimization is going to be critical and thank you so much for sharing this paper as this also validates what we are trying to do. So if we managed to keep the baseline below 40% through context optimization then coordination might actually work very well and helps at scaling agentic systems.

I agree on measuring and it is planned especially once we integrate the context optimization. I think the value of context optimization will go beyond just avoiding compacting and reducing cost to providing more reliable agents.

nextaccountic 12/28/2025|

> License: BSL 1.1, free for personal and educational use, converts to MIT in 2030 (would love input on this, does it make sense to just go MIT?)

I LOLd at that. Things in AI space become obsolete much faster. I'd say just go with GPL or AGPL if you don't want proprietary software to be built on top of your code

bahaAbunojaim 12/28/2025|

Converted it to MIT so it is MIT now

nextaccountic 12/28/2025|||

That's quite nice of you!

bahaAbunojaim 12/30/2025||

Thanks!

More comments...