July 27, 2026

Supervisor 2.1

Since 2.0 shipped, every flag Supervisor made in Discord came with a 👍 and a 👎 button. Thousands of you pressed them. Those votes are the single most valuable thing we have, because they are not our opinion about whether a decision was right, they are yours, on your own server, about your own members.

We read them. The pattern was uncomfortable and completely fair: Supervisor was too quick to flag. Not missing harmful content, the opposite. It was firing on messages that were obviously fine to a human, and every one of those costs a moderator time and makes the bot feel like a liability.

2.1 is a full retrain of all three models built directly on that feedback.

What changed

The headline number, on our internal evaluation set, comparing the Arbiter model in 2.0 against the Arbiter model in 2.1. Both are the same base architecture, so this is a like for like comparison:

2.02.1
Average F1 across 16 labels0.7940.941

Every one of the sixteen labels improved. The largest gains landed on the ones that were weakest before:

F1 score per label, Supervisor 2.0 versus 2.1 (Arbiter) Supervisor 2.0 Supervisor 2.1 0.00 0.25 0.50 0.75 1.00 Medical/Injury 0.98 +0.28 Violence 0.89 +0.24 Sexual (Unlawful) 0.93 +0.21 Promotional 0.95 +0.19 Sensitive 0.95 +0.19 Hate/Racism 0.95 +0.19 Spam 0.96 +0.19 Harassment 0.89 +0.14 Illegal 0.90 +0.13 Insult 0.93 +0.11 Self-harm 0.97 +0.11 Toxicity 0.93 +0.11 Scam/Incoherent 0.96 +0.10 Profanity 0.96 +0.06 Sexual (Explicit) 0.96 +0.05 Sexual 0.95 +0.04
F1 score per label on our evaluation set, Supervisor 2.0 against 2.1, Arbiter model. Sorted by improvement.

The same numbers as a table:

Label2.02.1Change
Medical/Injury0.6980.980+0.282
Violence0.6490.887+0.238
Sexual (Unlawful)0.7170.929+0.213
Promotional0.7520.946+0.194
Sensitive0.7610.950+0.190
Hate/Racism0.7650.954+0.189
Spam0.7720.959+0.188
Harassment0.7410.885+0.145
Illegal0.7650.897+0.132
Insult0.8150.929+0.114
Self-harm0.8570.970+0.113
Toxicity0.8180.927+0.109
Scam/Incoherent0.8570.957+0.100
Profanity0.9040.964+0.061
Sexual (Explicit)0.9130.960+0.047
Sexual0.9140.955+0.041

Violence and Medical/Injury were the two worst labels in 2.0, both under 0.70. They are now the two best.

The false positive problem specifically

F1 rewards catching things, so it is not the number to look at if what you care about is the bot leaving innocent people alone. For that, the metric is how well the model recognises a message as safe.

Across the sixteen labels, that score went from 0.975 to 0.986. That sounds like a small move until you look at it as errors rather than successes: the rate of getting a harmless message wrong dropped from 2.5 percent to 1.4 percent.

Error rate on harmless messages (lower is better) Mistakes on harmless messages, lower is better Supervisor 2.0 2.5% Supervisor 2.1 1.4%
Share of harmless messages the model gets wrong, averaged across all sixteen labels.

That is 44 percent fewer mistakes on safe content.

That is a real improvement and it is the one the feedback was asking for. It is not a claim that false positives are solved. If Supervisor flags something it should not have, press 👎. That is exactly the signal that produced this release, and it is what will produce the next one.

Short messages

The most common complaint we received was single words getting flagged. Messages like "join" or "wait" were being classified as spam, which is a bad look in any server where people talk normally.

This is a known weakness of models trained mostly on longer text: a very short message carries almost no context, and the model reaches for whatever weak signal it can find. 2.1 was trained with substantially more short-message data, and the Spam label specifically went from 0.772 to 0.959.

The tiers

All three models were retrained. Observer and Sentinel now use smaller, faster base models than 2.0 did, which is what lets them stay cheap, so their numbers are not directly comparable to the Arbiter figures above. Average F1 on the evaluation set:

TierAverage F1
Observer0.874
Sentinel0.891
Arbiter0.941

If you are on Observer or Sentinel and you moderate a busy server, Arbiter is where the accuracy is.

Being straight about what this is

These numbers come from our own evaluation set, which we built and label ourselves. That is the right tool for tracking whether a retrain improved things, because it is the same yardstick across both versions. It is not an independent audit, and you should read any vendor quoting their own benchmark, including us, with that in mind.

What we can point at that is not self-scored is the feedback itself. The 👍 and 👎 votes are yours, they are recorded, and they are what we are steering by. That is a slower signal than a benchmark, and a more honest one.

Rolling out

2.1 is live now for every tier, on every plan, at no extra cost. There is nothing to update. If Supervisor is already in your server, it is already running the new models.

Try it on something that used to get wrongly flagged in the live demo, or add Supervisor to your server.

And keep pressing the buttons.

Back to all posts

Stop Harmful Content Today

Join many communities already protected by Supervisor's AI moderation