Just don't expect to end up with a finished project; it's more like a first draft. Once it's there, it's much easier to determine what it is you actually want, since you can directly experience what works and what should be changed.
One important caveat is that I do not work with agents; each step goes through a fairly rigid manual review phase.
For complex features there can be 10 or more questions but I have a very strong sense of understanding the changes about to be made and Claude is very good at following all the decisions exactly. It's like working with an engineer who is both excellent at soliciting requirements and fast at implementation.
[0]https://github.com/mattpocock/skills/blob/main/docs/producti...
https://github.com/mattpocock/skills/blob/main/skills/produc...
That writing style is borderline incomprehensible.
I am still actively working on thesis : a self-directed plan mode to generate artifacts that can go through multiple evaluations of interactive interrogation is valuable
https://github.com/samelie/claude-plugin-pnpm/blob/main/skil...
It can still be used in ways that I personally consider correct, but I think the parts I personally consider incorrect are so inherently alluring that I find plan mode to be an overall net negative for software development. I celebrate its apparently impending default removal (at least Claude Code and OpenCode are openly stating that they think it's time for it to go).
Building an issue tracker that addresses this. It can replay the workflow after the fact like a movie, and pin down the parts that require human input via tagging and inline diffs in the tickets. It is git-native, lives in your repo alongside the code doesn’t require any external service.
Qwen 3.8 27b is the supervisor
Qwen 3.5 4b are the 6-15 minions it controls
Gemma 4 e4b is the validator for the supervisor.
A plan means it preps all work for the agents up front, tests that evals work, makes sure the dev environment is right for each agent, then finds and fixes each before the distributed tasks even begin.
What I thought would take minutes took hours as a supervisor or one agent did the prep / pre flight work.
My solution so far has been to drop all but basic setup and force the supervisor to ask before every op - if this is not the design choices, can this be run in parallel? If so, hand it off NOW.
I'm still iterating this workflow, but less setup for all the minions plus handing them work that may be incomplete/ broken is caught and fixed by the minion and its own qa gates.
This can mean a number of minions end up replicating the same fixes, but in general the time cost of that is small Vs the supervisor working in parallel instead of too sequentially.
This seems to be a good example because things like the menu, high score boards etc are common, but the games are distinct. Then there's the artwork which requires decisions on look, and for coordinating.
The Qwen 4B model is multimodal so part of the AC is to view the output - I've a robust anti AI-look QA chain for that I've been using elsewhere, e.g. no floating parts, consistency, obvious missing fingers etc etc.
The longer term plan is to do some llama.cpp refactors specifically for some target hardware I have and implementing slightly different novel architectures I'd like to try (one I did already targeted CPU inference, which I did using 3 agents with specific roles; main planner, QA for planner, and benchmarking/environment handling)
The implementation was 85% of the speed of the original maxed out on my hardware but performance scaled with CPU core count whereas the original implementation plateaued. Unfortunately the break even mark seemed to be around 30 - non HT - threads.
I suppose I should look at that one again, since the increase in cores did not linearly drop off performance e.g. due to memory contention.
> Qwen 3.8 27b is the supervisor
>
> Qwen 3.5 4b are the 6-15 minions it controls
>
> Gemma 4 e4b is the validator for the supervisor.
I just use Opus 5.5 and don't think about it?- Cold starts impact, context length issues, task lifecycle management
- Inefficiencies in delegation, necessitating workflow patterns for small projects (big AIs hide this problem until you scale and they hit the same issues).
- Limits of the AI would be harder to find or notice (e.g. where time - and cost - is being spent needlessly).
Some people are resource constrained.
For example, it’s a great hook in the process for agentic review.
Get a second agent to look at what will be implemented and check it for inconsistencies, check it against whatever decisions were made or provided previously in the chat, or against whatever technical rules you’ve written out for your project, before going ahead. It surfaces a large number of opportunities for refinement, and generally pushes the output closer to the direction you’re looking for.