The Problem with LLM Memory
Large Language Models have a fundamental limitation: finite context windows. Even with models supporting 100K+ tokens, autonomous agents that run for hours or days inevitably face memory degradation:
- Context overflow — important early information gets pushed out
- Retrieval inefficiency — finding relevant past context becomes a needle-in-a-haystack problem
- Memory decay — the model's ability to reason about distant information degrades
Our Solution: FAMM
FAMM (Future-Aware Adaptive Memory Management) is a framework we designed to address these challenges. It was published on Zenodo with DOI 10.5281/zenodo.21168000.
Core Architecture
Incoming Context
↓
┌─────────────────────────────┐
│ Adaptive Memory Organizer │
│ (categorize + prioritize) │
└─────────────┬───────────────┘
↓
┌─────────────────────────────┐
│ Future-Aware Prioritizer │
│ (predict future relevance) │
└─────────────┬───────────────┘
↓
┌─────────────────────────────┐
│ Intelligent Retriever │
│ (multi-strategy retrieval) │
└─────────────┬───────────────┘
↓
┌─────────────────────────────┐
│ Dynamic Memory Optimizer │
│ (compress + consolidate) │
└─────────────────────────────┘Key Innovation: Future-Aware Context Prioritization
Most memory management approaches for LLM agents are backward-looking — they store what happened and retrieve based on similarity to the current query. FAMM is different because it's forward-looking:
- It predicts which memories will be needed in the future based on the agent's current task trajectory
- It proactively elevates relevant memories before they're needed
- It compresses or discards information that's unlikely to be useful
Think of it like a chess player: you don't just remember past moves, you think about which past patterns are relevant to your next moves.
Memory Organization Tiers
FAMM organizes memory into adaptive tiers:
- 1.Active Memory — currently relevant context (always in the prompt)
- 2.Warm Memory — recently accessed, high future-relevance score
- 3.Cold Memory — compressed summaries of older interactions
- 4.Archive — highly compressed, retrievable only on specific queries
Intelligent Retrieval
Instead of single-strategy retrieval (e.g., just semantic similarity), FAMM uses multi-strategy retrieval:
- Semantic search — find contextually similar memories
- Temporal recency — prefer recent, relevant information
- Task dependency — retrieve memories that are prerequisites for current tasks
- Predictive loading — pre-fetch memories likely needed in the next 2-3 steps
Experimental Results
We evaluated FAMM against baseline approaches (sliding window, naive RAG, and fixed-buffer memory):
| Approach | Task Continuity | Retrieval Accuracy | Memory Efficiency |
|---|---|---|---|
| Sliding Window | Low | N/A | High |
| Naive RAG | Medium | 72% | Medium |
| Fixed Buffer | Medium | 68% | Low |
| **FAMM** | **High** | **89%** | **High** |
FAMM showed particular strength in:
- Multi-session tasks — maintaining context across conversation boundaries
- Complex reasoning chains — keeping track of intermediate results
- Long-term planning — remembering goals and constraints from early in the interaction
Why This Matters
As LLM agents move from single-turn chatbots to autonomous systems that operate for hours, days, or continuously, memory management becomes the bottleneck. FAMM is a step toward agents that can:
- Run customer support systems that remember user history across weeks
- Manage software projects with full context of every decision
- Conduct multi-day research with coherent, evolving understanding
Read the Full Paper
By Suyash Vakhariya and Asmita Ipper
