Memory in AI agents
Useful memory
Case state in the system of record, config versions, prior human decisions with reasons.
Retrieval over approved knowledge—not every chat the model ever had.
Dangerous memory
Unbounded chat history that re-injects stale or sensitive content into new runs.
"Learning" from production without review—quietly training on bad overrides.
Case state vs model memory
Case state belongs in your database with ids and status. Model memory is optional retrieval over approved knowledge.
If product language confuses the two, you will mis-design retention and access control.
Forgetting is a feature
TTL sensitive context. Do not keep chat forever by default. Align retention with compliance policy, not with convenience.
In practice
Map the workflow on a whiteboard before you open a framework: inputs, systems of record, humans, and irreversible writes. If that map is fuzzy, the agent will encode the fuzz.
Pick ten to fifty real historical cases as an eval set. Include the ugly ones. Run the agent offline against them until critical fields and hard rules are acceptable. Only then connect write tools.
Ship with a pause switch, a human queue, and a weekly review of override reasons. Promote repeated overrides into rules. That loop is how production systems improve—not another prompt brainstorm.
Common failure modes
- Treating a demo on clean samples as readiness for production volume.
- One shared service account with broad write access across systems.
- No owner for the exception queue, so failures pile up as noise.
- Changing prompts and models without regression gates on real cases.
- Measuring only model latency or thumbs-up, not completed-case cost and audit completeness.
What good looks like after ninety days
The first workflow is boring: stable override rate, known failure modes, operators who trust the queue. Config changes go through review. Traces answer "what happened to this case?" without archaeology.
At that point you can add a second document type or a second agent role. Expanding before the first path is boring is how programs stall with five half-built pilots.
Frequently Asked Questions
Do agents need vector memory?
Sometimes for retrieval. They always need explicit case state. Do not confuse the two.
How long should we retain agent logs?
Follow your compliance retention policy—design it before go-live.