AN OPEN BENCHMARK FOR AI MEMORY · OCT 6, 2026Lost Thread
27 questions about your own week. Same models, with and without DayDream.
Helpful answer rate
n = 27 per setup
ChatGPTGPT-6.1
ClaudeOpus 5.5
All results
| ChatGPT | ChatGPT + DayDream | Claude | Claude + DayDream |
| Helpful answer rate1n = 27 | 18.5%8.2–36.7 | 51.9%34.0–69.3 | 3.7%0.7–18.3 | 66.7%47.8–81.4 |
|---|
| Accuracy2n = 27 | 27.8% | 61.1% | 13.0% | 74.1% |
|---|
| Head-to-head win rate3n = 27 | 14.8% | 59.3% | 11.1% | 88.9% |
|---|
| No made-up facts4n = 27 | 96.3% | 92.6% | 100.0% | 92.6% |
|---|
| No overclaims4n = 27 | 100.0% | 92.6% | 100.0% | 96.3% |
|---|
Two of the 27
ChatGPT“What was the Airbnb reservation number for the Fredericksburg cabin?”
Own memoryHelpful 1/4
I don’t have the Airbnb reservation number in our chat history. You can find it in your Airbnb confirmation email or under Trips → your Fredericksburg booking → Reservation details.
With DayDreamHelpful 4/4
The reservation number was HB-88231, shown on the Airbnb reservation page you viewed Friday, Oct. 2 at 6:07 PM.
Claude“What did we end up deciding for the consumer group naming convention?”
Own memoryHelpful 1/4
I don't have the specifics of that decision. My notes from our past conversations say we covered consumer groups as part of the Kafka migration, but not which naming convention you settled on, and I don't want to guess and have you bake the wrong pattern into the migration.… (answer continues, 1,278 characters)
With DayDreamHelpful 4/4
The convention you decided on is <app>-<domain>-<env>-v<version>, for example order-sync-prod-v2.
You wrote that into the Kafka Event Bridge Runbook on Confluence on Wed Sep 30 at 12:21 PM: "Consumer group naming convention: <app>-<domain>-<env>-v<version>, e.g. order-sync-prod-v2. Team voted 4-2 on 6/5."
You opened the "Consumer Group Naming" Google Doc again on Fri Oct 2 and yesterday at 10:52 AM and 12:47 PM, and Slack with Marcus and #integrations was open at the same times. I can't see what you typed during those visits, so I can't tell whether anything changed. If someone raised it again in the last few days, check the latest version of that doc.
How it was run
- Three made-up work weeks, 27 questions, locked before any model ran.
- Each model answered with its own saved memories, then again with DayDream added.
- Three judges graded every answer without knowing which setup wrote it.
- Every head-to-head was judged twice, with the sides swapped.
- Everything is open: questions, answers, grades and judge instructions.
- Graded 3 or 4 out of 4 for meeting the person’s need. Median of 3 Claude Sonnet 5.5 judges. Small numbers are 95% intervals.
- Mean correctness grade as a share of the maximum.
- Share of questions where the judges preferred that answer. ChatGPT tied on 25.9%.
- A made-up fact is a claim the week doesn’t support. An overclaim states something DayDream can’t know, like whether a message was sent.
ChatGPT is GPT-6.1 through Codex; Claude is Opus 5.5 through Claude Code. With DayDream available, ChatGPT used it on 16 of 27 questions and Claude on 26. Longer answers didn’t score higher (r = −0.08). People and names are invented.
Try it on your own week.
Free & open source · Apple silicon · macOS 15+