Top
Best
New

Posted by Anon84 2 days ago

Building an Advanced Agentic Harness(data4sci.com)
129 points | 42 comments
hanneshdc 2 days ago|
Any benchmarks showing if this actually improves problem solving? Or reduces errors?

The idea is cool, but from own experience in harness engineering, lots of cool sounding ideas can have a negative impact on performance due to emergent and confounding effects.

So I'm a bit skeptical!

Supermancho 2 days ago|
The TUI and Session manager are straightforward enough. PI is serial by default but can become more DAG-like, where Codex is designed specifically for DAG and this is going more the Codex direction.

I'm much more interested in the memory model and why. As far as I can tell, it's "a vector Db" and not much more is said. Nothing about working memory or procedural memory (there are lots of ways to classify it, https://www.youtube.com/watch?v=BacJ6sEhqMo), but I was disappointed with how "advanced" it seems.

troupo 2 days ago||
Question: Any benchmarks showing if this actually improves problem solving? Or reduces errors?

"Answer": a word soup that in no way, shape, or form addresses the question, but does sound jargony and vague enough to be an LLM.

ilaksh 2 days ago||
So if I understand correctly, one of the agents is creating the workflow as a DAG dynamically for each new job it's given? That's the most interesting part to me. The rest I was already pretty familiar with.

So maybe that's a big chunk of what you need for an 'AI Company': an agent that manages the goals and hierarchies. Although of course the DAG and agent hierarchy is not quite the same thing. But maybe the workflows and subworkflows are what matter.

Anon84 2 days ago|
Yes, there's a team of agents, each with a different role.
abdullahkhalids 2 days ago||
A harness [1] was developed by Terrence Tao and some collaborators to prove mathematical results. It has since then been used by others with positive effect. Can someone critique the structure of this harness? I don't know anything about this stuff.

[1] https://github.com/1stproof/batch-2/tree/main/batch-2-submis...

cyanydeez 2 days ago|
it's not actually proving it though? It's more like stringing it together. A person or LEAN has to actually provde something. I've yet to see anything other than AI-slop produces simulcra of proofs. If it were proving something it'd be <insert mathematician> validates AI proof.
abdullahkhalids 2 days ago|||
I don't you what you mean. People have used the above harness (or similar) to prove significant results. See this recent paper [1], which claims

> The human authors take full responsibility for the claims and proofs contained in this paper, and have carefully refined and verified them. The construction and main ideas of the proof were generated entirely by Codex using GPT 5.6 Sol Ultra, using harness ideas generated by the authors based on the UCLA Moonshot Harness [ZHC+26] and [Ope26].

[1] https://arxiv.org/pdf/2607.21551 (Statement on AI usage is at the bottom of page 3).

cyanydeez 2 days ago||
you keep using the term "used the model" or whatever.

A model is non-deterministic. People prove things, LLM string together a bunch of words and do symbol shunting.

Ensure you understand what symbol shunting is before you make claims. https://ell.stackexchange.com/questions/76400/what-does-one-...

Real break throughs come from integral mathematics and not just a few reorderings. I've no doubt these are talented people recognizing output as useful; however, every time I see these links presented it's never from the "Prominent mathematician verifies AI proof"

Don't put the cart before the horse if you want people to think LLMs are cracking math problems in real terms.

antonvs 2 days ago|||
Stringing what together? A sequence of logical implications? The word for that is "proof".
cyanydeez 2 days ago||
You mis understand, which is why these articles are so empty. I could find some math starved journal and poblish a bunch of logical implications, but that doesn't mean the logic is sound.

A collection of logical implications <> proof.

antonvs 2 days ago||
A sequence of logical implications that each follow from previous ones is exactly what a proof is. Did I really need to spell that out?
DerrickDevo1 2 days ago||
A good tutorial. Generally speaking, the harness is the environment layer between a language model and its task, such as the action set it can call, the state, the context it can see and the memory etc.

However, currently the bigger question comes to my experience during harness is actually not where we use LLM in the system, but where we do NOT use LLM in the system. And the validation of the results becomes more and more important. Any thoughts on this?

bryan0 2 days ago||
I (like presumably many others) have built something similar. My main difference though is that the critics operate on each stage of development before it can move onto the next. The stages are defined by deliverable artifacts: issue, plan, pull request. Critics must approve each artifact before you can move onto the next. So the process is defined by a DAG which defines how to transition successfully from one artifact to the next. It's been fun to work on and I would like to open source it soon, but I assume many others are working on similar systems.
budududuroiu 2 days ago||
> The plan is a graph

I much prefer giving the LLM a REPL loop, and injecting all the tools as functions inside the REPL loop.

That means that the LLM isn't constrained to writing a DAG, it can write code that loops, exits early, etc.

Axsuul 2 days ago|
Do you have an example?
lmeyerov 2 days ago||
A graph is a fancy way of saying a few async await calls, which a REPL can do

(We added the same to louie.ai, not complicated)

jumploops 2 days ago||
Contrary to the title and intro, this appears to be an agentic _workflow_ builder/runner, not an advanced “agent harness”

A few things:

- they note: “nothing in this post proves it actually works in most cases”

- the DAG sounds good, but LLMs often split tasks into smaller pieces than they need to, which can cause them to lose the forest for the trees

- the forced JSON interplay, in my experience, causes even gpt-5.6-sol to lose a few “IQ points”

For anyone reading this, this tutorial is much more reminiscent of how folks were building “agents” pre-Claude Code.

tl;dr the “orchestrator” here is just a software loop, and the LLM prompts restrict flexibility of the planner/workers

Axsuul 2 days ago||
Anyone else have related reading that touches on this? I'm building my own custom harness and want to start implementing loop support, etc. But I also want to build some sort of framework so that it's dynamic (e.g. this needs to run x number of iterations, while planning needs to run y number of iterations).
zarldev 2 days ago||
OpenAI and Anthropic both have blog posts about programmatic tool use. Tbh I settled in STARLARK as the intermediary language it's a python dialect so most llms can use it with little help.https://github.com/zarldev/zarlmono/blob/main/zkit%2Fagent%2... it's a wrapper around the calls of the other tool set.
metadat 2 days ago|||
Check out https://github.com/strongdm/attractor along with the community implementations at https://factory.strongdm.ai/products/attractor#community.
hagen8 2 days ago|||
Check out academic papers about:

1. Hierarchical skills, workflow, skill learning 2. Meta Harness, self-learning harnesses 3. Trace/trajectory representation 4. Common agentic benchmarks

But first more basic things like 5. Blog posts form anthropic 6. How Claude Code/PI/ Hermes!! agent works 7. Agent sessions/ Forking/ Hooks

alansaber 2 days ago||
Sounds like you want something similar to /goal mode in Codex.
Anon84 2 days ago||
And the associated GitHub repo: https://github.com/DataForScience/LLMs
andai 2 days ago|
Do big models benefit from Conway's Law, or only small ones?
More comments...