Wiki / Core Features / AI Responses / AI Memory

AI Memory

Updated by Maxime_48 · 1 week ago · 132 views

AI Memory lets YAWBDB store useful context for a server so AI responses can become more consistent over time.

Two things write to it: you, from the panel, and the bot, which notes what it learns from conversations on Discord. The memory system is managed from the panel, and nothing in it is ever removed automatically.

The AI Memory page listing live and retired facts for a server, with a memory type breakdown card above the table


🧠 What memory is used for

Memory can help the AI remember information such as:

  • Server lore
  • Community facts
  • Preferred tone
  • Recurring jokes
  • User or role context
  • Server-specific terminology
  • Important rules or reminders

The goal is to make AI responses feel more aware of the server.


📒 Memory page

The Memory page gives authorized users a full view of stored memory entries.

From this page, you can:

  • View memories by server
  • Search by key or content
  • Add new memories manually
  • Edit existing memories
  • Delete outdated memories
  • Review confidence values

🔎 Filtering and pagination

Memory can be filtered by server and searched instantly.

Large memory sets are paginated so the page stays usable.


✍️ Manual memory entries

You can add memory entries manually when you want the AI to know something specific.

Examples:

  • server_theme: This server is about retro gaming.
  • bot_tone: Replies should be friendly and concise.
  • important_rule: Do not give support in public channels; direct users to tickets.

Use short, clear entries. Avoid dumping huge paragraphs unless necessary.


✏️ Editing memory

You can edit:

  • Key
  • Content
  • Confidence

Edit a memory when it is still useful but needs correction.

Delete it when it is outdated or no longer relevant.


♻️ Nothing is forgotten behind your back

Entries are never removed automatically. There is no expiry and no clean-up job: a fact stays until the bot itself decides it no longer holds, or until someone removes it from this page.

What you control is how much of it reaches the AI. Max memory facts injected, on the AI Responses settings page, sets how many facts are sent as context — from 1 to 1000, 20 by default. Raising it gives the AI more to work with; lowering it keeps answers tighter. Nothing is deleted either way, so you can move it up and down freely.

Age is not a reason to forget. This server speaks French is still true two years on, and nothing will restate it.

How the bot decides

Each fact sent to the model carries its own number, shown as [#12] in the context it receives. When the bot concludes that something no longer holds — a rule that changed, a preference someone reversed — it answers with that number, and the entry stops being used. Whether it is deleted or kept for reading is what the memory type decides, below.

It can only remove a fact it was actually shown in that exchange. If it names a number that was not in front of it, nothing happens: a guessed number cannot reach a fact from another conversation, even a real one on the same server.

This is why Max memory facts injected matters twice over. A fact that never reaches the model can never be re-examined by it — it stays until someone removes it from this page. Under the searchable archive that stops being a dead end: a fact the bot looks up is a fact it can also correct.


🧭 Choosing a memory type

Memory type, on the AI Responses settings page, decides how the bot stores what it knows and how it gets it back. It is not a dropdown: click it and the six types open over the page, side by side, each with what it does, what it changes, what switching to it would move, and where to read more.

Classic is what runs today and it is the default. A server that never opens this field keeps behaving exactly as it always has: a flat list of facts, the most recently updated ones injected into every answer, nothing removed on its own.

The other five each add one thing on top of the previous one, and they are being built in order:

  • Journal — available now. A fact that stops being true is retired rather than erased, so you can still read what it used to say. See the next section.
  • Searchable archive — available now. The bot may look something up instead of only using the facts it was handed, so what sits past your injected limit stops being invisible. See the section after next.
  • Ranked memory — available now. The slots go to the most useful facts rather than merely the most recent ones.
  • Consolidated memory — available now. A background pass brings duplicates together and rewrites the summary while the server is quiet, and proposes every one of them for you to approve.
  • Semantic memory — matching on meaning rather than on words, so "he speaks French" finds "the server language is French". The strongest of the six, and the only one with a bill attached.

The one type that is not ready yet is marked in the dialog and cannot be selected. Whatever the type, the same two rules hold: they all read the same memory, and none of them deletes anything on its own — so changing your mind, and changing back, never costs you a fact.

With one server picked, the memory page tells you what a change would actually move, in counts, before you make it.


📓 Journal — nothing is erased, only retired

Under Journal, the bot no longer deletes a fact it decides is out of date. The fact's period of validity simply ends: it leaves the context the AI is given, and it stays on this page with the date it stopped and, when there is one, the fact that took its place.

It also gains the move that goes with it. Instead of only being able to drop [#12], the bot can say "this one is no longer true, here is what replaces it" — and it is told, in so many words, that the old wording survives. A correction it believes is destructive is a correction it hesitates to make.

Reading the history

Three tabs appear above the table once something has been retired:

  • In use — what the bot can actually see. This is the default, and it is the page's real answer to "what does the bot know?".
  • Retired — what used to be true. Struck through, with the date, the reason if the bot gave one, and the replacement.
  • All — both, in one list.

Undoing a wrong call

The bot gets it wrong sometimes. Restore, on a retired row, reopens it in one click — that is the whole reason the row was kept instead of deleted. The replacement, if there was one, is deliberately left alone: two facts can be true at once, and which of the two goes is your call, not the panel's.

Delete still deletes, on a retired row like on any other. A person removing a fact from this page has decided; nothing second-guesses that.

Switching, and switching back

Turning Journal on moves nothing. Every fact you already have counts as still true — which is what it already was — and keeps the date it was created.

Going back to Classic leaves whatever was retired in the meantime exactly where it is: still on this page, still not seen by the bot. If you want it all back in front of the AI, the memory page offers Restore all retired facts for that one server, and writes it to your audit log. It is never done for you: what the journal retired, it retired for a reason.


🔎 Searchable archive — the bot can look things up

Classic memory has one real limit, and it is not the number of facts you may store. It is that everything past Max memory facts injected is invisible to the AI, however true it is. Raising the limit is the only lever, and it makes every single message more expensive.

Searchable archive gives the bot the other lever. The recent facts are handed to it as usual, and on top of that it may ask for one by name — "search the memory for: Paul's birthday". What it finds is given to it before it answers.

What it costs

Nothing on the messages where the bot does not look anything up, which is most of them. When it does look something up, it costs exactly one extra call to the AI — the question is asked again with the facts in hand. Never two: a bot that asks to search again in its second answer is ignored.

The search itself is full text, run by the database. No embeddings, no vectors, nothing computed per message, and no third-party service involved.

What it can reach

Every fact of that server that is still in use, in both scopes — facts about the server and facts about its members. That is deliberate: "Paul's birthday" is a fact about someone who is not the one talking, and a search that could only reach the speaker's own facts would miss most of the questions worth asking.

What it cannot reach: anything retired under a journal — closing a fact's window is precisely the decision not to hand it back — and anything at all in a channel where memory reading is not allowed. A search is a read, and it obeys the same Memory read channels setting.

A fact the bot found by searching is also one it may correct, on the same terms as the ones it was handed.

Before you switch: the dry run

With one server picked, the memory page tells you what the archive would cost you, in numbers taken from your own data: how many facts there are to index, how much text that is, and how long a real search over them just took.

Turning it on moves nothing and rewrites nothing. The index is built over the text that is already stored, and it is there for every server whether or not they search — so switching, and switching back, is instant either way.


🏅 Ranked memory — the slots go to the facts worth them

Classic memory hands the model the last N facts modified. Said another way: a rule your server has lived by for a year loses its place to yesterday's small talk, because the small talk was touched more recently.

Ranked memory scores every candidate fact on three things and hands over the best N instead:

  • How recent it is.
  • How important it is — scored by the model at the moment it writes the fact, which is the only moment the reason for storing it is still in front of it.
  • How much it has to do with what is being said right now.

It is a ranking, not a filter. Nothing is set aside for good, and a fact that misses the cut on one message is a candidate again on the next.

The weights are yours

Three settings on the AI Responses page — recency, importance, relevance — decide how much each leg counts, per server. A server whose memory is mostly standing rules will want importance to weigh more than freshness; a busy social server will want the opposite.

The confidence already stored with every fact is read too. It acts as a multiplier, and only ever downwards: a fact the bot was unsure about does not outrank one it was certain of.

Switching costs nothing

There is no migration and no bill. A fact nobody has scored yet ranks in the middle rather than last, so the ranking works from the first message on what you already have, and importance fills in on its own as the bot rewrites facts.

The dry run says how many of your facts already carry a score, and shows which ones would gain a slot and which would lose one.


🧩 Consolidated memory — the tidying happens while the server is quiet

A memory that grows ends up saying the same thing twice under two names: language on one side, lang on the other. Until now the only way to deal with that was by hand.

Consolidated memory adds a background pass over the memory of a server that has gone quiet. It never runs in the middle of an answer, and — this is the important part — it applies nothing on its own.

It proposes, you decide

The pass writes proposals. They appear on the memory page one by one, the fact kept next to the fact it would replace, and each one has three answers:

  • Merge — take the pass's suggestion.
  • Merge the other way round — keep the other fact instead. On two facts that disagree, you are the one who knows which is still true.
  • Keep both — decline. The pass will not ask again unless one of the two facts is rewritten afterwards.

Nothing changes until one of those is clicked.

A yes never erases

A merge is a replacement in the journal sense. The fact that steps aside closes its window pointing at the one that stayed, it shows up struck through on this page with its date and its reason, and Restore brings it back. A merge is undone exactly like any other retirement.

Contradictions are never merged. Two facts that say opposite things are flagged for a person to settle. Nothing in the pass knows which of them is still true.

The summary gets rewritten

The summary is the one memory row the bot reads whole, every time — so it is also the row that goes stale first. It was written once, from what was known then.

The pass drafts what it would say today from the server-wide facts, and the page puts the two side by side: what your bot reads today, and what it would read instead. Same rule as a merge — nothing replaces it without a yes, and the previous wording is retired rather than deleted.

Its own model

This pass is not racing a Discord reply, so it can run on a stronger model, or a cheaper one, different from the one that answers in the channel — with its own endpoint and its own key, like every model field in the panel.

It also works with no model at all. Without one, the pass still finds and proposes every pair; a merge simply keeps the winning fact word for word instead of rewriting the two into one sentence. The model only ever writes the suggested wording and the summary draft.

When it runs

On its own schedule, and it refuses itself whenever a pass would not pay: a memory too small to hold duplicates, nothing changed since the last pass, less than a day since the last one, or a proposal already waiting for an answer. A server that is never quiet gets its pass anyway after a week, because that is exactly the server whose memory fills with duplicates.

Before you switch: the dry run

The card tells you, on your own data, how many facts are stored, how many duplicates and how many contradictions a pass would find, with five real examples taken from your server.

Turning it on moves nothing. The first pass proposes; you decide.


🧠 Semantic memory — meaning instead of words

Every type before this one answers the same question underneath: do these two texts share words? For "he speaks French" and "the server language is French", the answer is no — one word in common, and it is not the useful one.

Semantic memory turns each fact into an embedding: a direction rather than a string. Two facts that mean the same thing point the same way, whatever words they were written with, and the closeness between them is a number the panel can compare.

What actually changes

Two things, and the second one is the one you feel every day.

The search stops missing things. When the bot goes looking for a fact it was not handed, it stops depending on the member having used the same word as whoever wrote the fact down.

The ranking stops depending on the words of the message. Under Ranked, a fact scores relevance from where it landed in the search results. Under Semantic it scores on how close it actually is to what is being said — the difference between "this fact was found" and "this fact is about this".

What it costs

This is the one type that sends your data somewhere on a schedule, so the cost is worth stating plainly:

  • One embedding call per fact written, once. A fact that is not reworded is never embedded twice.
  • One call per message long enough to be ranked against. Under twelve characters, nothing is sent — "lol" has nothing to match on.
  • The comparison itself happens in the panel, not in the database, which is why the type is comfortable up to a few thousand facts on a server and not far beyond.

It needs an embedding model, named on the same endpoint and key as the rest of your AI settings.

It falls back on its own

No model named, an endpoint that refuses, a catch-up pass that has not run yet, or an embedding model you have just changed — in every one of those cases the search goes back to matching words instead of failing. The answer says which of the two actually ran, so a silent fallback stays visible.

Changing your embedding model re-embeds everything, and until that pass is done the server searches by words. Nothing breaks in the meantime.

Before you switch: the dry run

The card tells you how many calls the catch-up pass would make, over how much text, and which of your facts the vectors already consider the same thing.

No price in currency. Putting an amount there would mean inventing a tariff for whichever endpoint you named, and a made-up number on a page about money is worse than no number at all.


🧹 Clearing memory

Deleting entries one at a time is right for a single wrong fact. For two other cases there is Clear memory, next to the server filter once you have picked one server:

  • Everything about this server — when the memory has drifted far enough that starting over beats twenty corrections.
  • Everything about one member — when someone has left, or asks to be forgotten. Enter their Discord user id.

It is immediate and cannot be undone, and the bot starts learning again from the next conversation. Every clear is written to your audit log under the name of whoever did it, with how many entries went.


🔑 Access control

Memory entries are visible to anyone with access to the team that owns the server.

Changing them is a different right: adding, editing, deleting and clearing all need permission to configure the server — the same one the feature settings need. Without it the list is still yours to read, and the Clear memory button is simply not there.

Do not store sensitive secrets, passwords, private tokens, or confidential information in AI memory.


✨ Memory and response quality

More memory is not always better.

Too many irrelevant facts can make the AI less focused.

Good memory should be:

  • Short
  • Accurate
  • Relevant
  • Up to date
  • Easy to understand

💡 Best practices

  • Review memory after major server changes.
  • Delete outdated facts.
  • Clear a member's entries when they ask to be forgotten, rather than hunting for their rows.
  • Keep important rules clear and explicit.
  • Do not store private user data unless you have a good reason and permission.
  • Use memory to improve context, not to replace moderation or staff judgment.