A long output is different. A system can write many words without making progress on the assigned task. An agent’s value comes from the results of its actions, not the length of its response.
The working loop
- Plan: decide what information or action is needed next.
- Act: use an allowed tool or produce a proposed change.
- Observe: inspect the real tool result rather than assuming success.
- Verify: compare the outcome with the task’s acceptance criteria.
- Continue or stop: choose the next step, request review, or finish.
This is a useful explanation of the workflow, not a specification of Argon’s internal architecture.
Why intermediate state matters
Connected steps depend on keeping track of what has already happened. A coding agent needs the current repository state and test results. A research agent needs its sources, open questions and conclusions. If that record is lost or inaccurate, later actions can repeat earlier work or rely on assumptions that no longer hold.
Save enough information to resume after an interruption and to let a reviewer reconstruct the important decisions. The record should include failed actions as well as successful ones.
How to evaluate an agent
Check the complete task rather than scoring the final explanation alone. Did the actual change pass the independent check? Did the research cite evidence that supports its claims? Did the system stay inside its permission and spending limits?
Define stopping conditions before the run. A useful agent can recognize a blocked task and leave a clear record for the person taking over. Read the safety guide for permissions and review boundaries, or apply the coding protocol to a repository task.