Policy
What the agent may never do without a person
The strongest limit on what an agent can do wrong is what it is permitted to do at all.
FlowFinds runs an agent that spends money, approaches suppliers, writes to customers and holds data on behalf of the people who use it. Most of our safety work is not about making the model better behaved. It is about reducing the set of things a badly behaved model could do.
Three limits are absolute. The organic campaign tooling holds no account on any platform and publishes nothing — it produces artifacts a person posts. The agent does not move money; spend is initiated through controls a person operates and against a balance a person can see. And an irreversible action requires a named human, which in an exclusive inventory covers rather more ground than it sounds: passing a product, claiming one, and releasing a hold are all irreversible.
There is a matching rule in the interface. Before any control that cannot be undone, the consequence is stated in advance. Passing a product tells you plainly that you will not see it again. Moving your hold tells you the previous one is released. A control whose consequence is discovered after the click is a trap.
There is also a list of things the agent is forbidden to say, which is enforced rather than encouraged: it may not assert a quoted figure as established, and it may not report an action as taken when no action was taken. Those two are scored on the benchmark, which is the point — a limit nobody measures is a preference.
Read next
The published limits on agent authority.
More from FlowFinds
- Newer: Money is credited on a signature, never on a success URL
- Older: A find is only as current as its most neglected axis
- Everything else is in the news index.
To fact-check anything above before you publish it, write to [email protected].