Evaluating Jev for content moderation
A comparison of Jev, GPT-5.6 Luna and GPT-6 Luna on four moderation datasets, including prompt trade-offs, historical cost and recorded request time.

A moderation check has two jobs: catch harmful requests and let harmless ones through. Blocking an ordinary question about a sensitive subject is a mistake too. Can Jev handle that decision well enough to be useful, and do it faster and cheaper?
The original tests compared Jev 1.13.0, GPT-5.6 Luna and GPT-6 Luna on four public datasets. Jev using the baseline prompt had the highest observed F1 on all four, although some differences were too uncertain to call. A later initial comparison with GPT-6 Luna through OpenAI Decisions also favoured Jev in this setup, with important differences in tuning and refusal handling.
What the datasets test
-
ToxicChat contains real chatbot requests labelled for toxicity, testing the messy language people actually use.
-
XSTest pairs harmless prompts that can look unsafe with genuinely unsafe ones, testing whether a model overreacts to wording.
-
Aegis 2 is NVIDIA’s content-safety dataset, covering a broad range of safety risks under a detailed labelling policy.
-
WildGuardTest includes ordinary and deliberately adversarial requests, testing harmful intent as well as misleading framing.
These aren’t interchangeable definitions of safety. An abusive message and a politely worded request for harmful help pose different problems.
How the models were tested
Each model classified message text as safe or unsafe, checked against dataset labels. The original tests didn’t assess generated answers or refusal behaviour. After exclusions, each Luna comparison scored 9,155 rows from 9,035 distinct API payloads. Exact duplicate payloads reused predictions while every eligible row kept its scoring weight. The later Decisions comparison and its classification refusals are discussed separately below.
Selected development and test inputs were separated by Unicode, case and whitespace-normalised exact text. A later similarity audit found 18 ToxicChat test rows, eight toxic, near development inputs. Excluding them in a post-test sensitivity check preserved the main conclusions. This doesn’t establish full semantic independence.
Jev returned a safety score, converted to a decision using a 0.54 cutoff. Here, Jev baseline prompt and Jev content-focused prompt distinguish two prompts for the same Jev 1.13.0 model, with the same cutoff.
The baseline prompt focuses on whether the message asks for harmful assistance. The content-focused prompt additionally considers harmful content in the message itself. This broadens the decision beyond the help requested, so the choice of prompt can change what the moderation check catches.
I chose the baseline prompt and cutoff using 200 WildGuard training examples, then checked them on an independent 200-example sample before the held-out tests. Both Luna models used the same safety criteria in a strict Boolean JSON prompt, returning true or false with reasoning disabled.
Both Jev prompts were frozen before the 1,923-row Aegis test and used the 0.54 cutoff without further tuning.
A false positive is a harmless message wrongly flagged. A false negative is a harmful message missed. F1 balances how often flagged messages really are unsafe with how many unsafe messages the model catches. It isn’t overall accuracy.
What the results show
| Dataset | Test rows | Jev baseline prompt | GPT-5.6 Luna | GPT-6 Luna |
|---|---|---|---|---|
| ToxicChat | 5,083 | 75.40 | 66.86 | 63.68 |
| XSTest | 450 | 93.40 | 88.27 | 91.46 |
| Aegis 2 | 1,923 | 80.27 | 79.12 | 77.58 |
| WildGuardTest | 1,699 | 88.30 | 83.54 | 85.17 |
F1 for unsafe or toxic messages, on a 0–100 scale. Higher is better. The Jev column uses the baseline prompt throughout.
Sensitivity checks used 2,000 paired bootstrap resamples of normalised-message groups. The Aegis comparison with GPT-5.6 and XSTest with GPT-6 were inconclusive. The other paired differences favoured Jev. These nominal intervals describe sampling uncertainty conditional on cached predictions. They weren’t adjusted for multiple comparisons and don’t capture prompt selection, repeated model runs or all shared-template dependence.
“Jev using the baseline prompt had the highest observed F1 on all four, although some differences were too uncertain to call.”
Figure 1. Baseline benchmark results, unchanged after the later GPT-6 prompt search. The charts’ original Jev results refer to the Jev baseline prompt.
On WildGuard, Jev caught 638 of 754 harmful requests, versus GPT-6’s 600, and falsely flagged 53 harmless messages versus 55. That’s 38 more harmful requests caught with two fewer false alarms. Jev still missed 116. GPT-6 improved on GPT-5.6 for XSTest and WildGuard, but fell behind on ToxicChat and Aegis.
What prompt changes showed
I selected the Jev content-focused prompt from three prompts tested on 1,000 ToxicChat training groups, then checked it on another 1,000 independent groups. The selected prompt was fixed before test inference and the 0.54 cutoff stayed unchanged.
| Dataset | Jev baseline prompt | Jev content-focused prompt |
|---|---|---|
| ToxicChat | 75.40 | 80.23 |
| XSTest | 93.40 | 90.72 |
| Aegis 2 | 80.27 | 80.77 |
| WildGuardTest | 88.30 | Not run |
F1 for the same Jev 1.13.0 model with two prompts, on a 0–100 scale. Higher is better.
The ToxicChat gain and XSTest decline were both supported by the uncertainty checks. The smaller Aegis change was inconclusive, and the content-focused prompt wasn’t run on WildGuardTest. This suggests that improving the fit to one safety definition can worsen the fit to another.
I kept the baseline in the main comparison because it has results on all four datasets and uses the same safety criteria carried into both Luna prompts. Choosing whichever Jev prompt scored best on each test set would be a different comparison. The content-focused prompt shows a trade-off, rather than a consistently better replacement.
“The ToxicChat gain and XSTest decline were both supported by the uncertainty checks.”
The GPT-6 Luna search tested three fixed prompts on 600 previously unused ToxicChat training groups. The original won with 69.70 F1, versus 68.75 and 66.67. On 600 separate confirmation groups, it scored 72.16 F1, wrongly flagging 2% of harmless messages. Following the pre-agreed protocol, I reused its historical four-dataset predictions: there was no full benchmark rerun. This small search found no improvement, rather than proving none is possible.
Cost and recorded request time
Historical estimated cost per thousand requests was $0.019–$0.022 for Jev, $0.060–$0.077 for GPT-5.6 and $0.029–$0.037 for GPT-6. These estimates use returned token usage and recorded rates, including cache writes.
Figure 2. Cost per 1,000 unique requests. Counts: ToxicChat 4,976, XSTest 450, WildGuard 1,699, and Aegis 1,915 for Jev and 1,910 for each Luna run.
Median recorded times for successful requests were 289–316 milliseconds for Jev, 707–720 for GPT-5.6 and 750–770 for GPT-6. Jev’s timer stopped when the HTTP call returned. Luna’s also included JSON parsing, response-file writing and label/cost processing. The models ran separately under different service conditions, and the saved timings can’t isolate local processing costs. This wasn’t a controlled inference-speed comparison.
Figure 3. Median recorded time for successful requests. Client queue and rate-limit waiting are excluded. No timing uncertainty intervals were saved.
An initial comparison with OpenAI Decisions
I also tested GPT-6 Luna through OpenAI’s Decisions endpoint. Two frozen prompts transferred the existing criteria: the original assistant-safety prompt, with its Chat output-format instruction removed, and a policy variant formatted from the Jev content-focused instructions. Each scored the same 9,155 rows from 9,035 distinct requests, using a 0.5 probability cutoff. Historical Jev and Chat predictions were reused. No Decisions prompt search or threshold fitting was performed, and no further tuning was conducted.
| Dataset | Jev baseline prompt | Decisions original | Decisions policy variant |
|---|---|---|---|
| ToxicChat | 75.40 | 50.00 | 62.88 |
| XSTest | 93.40 | 86.50 | 84.32 |
| Aegis 2 | 80.27 | 72.99 | 72.41 |
| WildGuardTest | 88.30 | 79.63 | 73.08 |
F1 on a 0–100 scale. Decisions scores include blocking requests when the endpoint refused to classify them.
Jev had higher observed F1 in all eight comparisons. The nominal 95% intervals from 2,000 paired resamples of normalised-message groups favoured Jev throughout, with the same sampling and multiple-comparison limitations described above. The policy variant improved Decisions on ToxicChat but reduced F1 elsewhere. Original Decisions caught 36.74% of toxic ToxicChat messages, compared with Jev’s 71.55%, while falsely flagging 0.78% of harmless messages, compared with 1.40%. Lower false-positive rates came with more missed harms.
Refusals need separate attention because they contain no classification probability. The primary analysis counted them as blocked deployment actions, so it measures the classifier plus that fallback. This departed from the earlier preflight plan to leave refusals unresolved, although the blocking rule was frozen before inference. Probability coverage was 99.64% of scored rows for the original prompt and 97.87% for the policy variant. On XSTest, the variant refused 9.56% of rows: allowing rather than blocking those requests reduced F1 from 84.32 to 72.78. An offline check excluded refusals and compared each configuration on exactly the same answered rows. Jev remained ahead in all eight comparisons, but those selected subsets differ between prompts and no new confidence intervals were calculated.
The run’s measured token usage gave a $0.6790464 base cost estimate, including the development pilot. A conservative ledger allowance totalled $1.0197696, including a 50% margin and a reservation for one failed attempt. Neither figure is an invoice-reconciled bill. Base estimates across datasets and prompts were $0.032–$0.043 per thousand unique test requests.
Median client request times were around 301 milliseconds, including network, server and local response processing. They used final valid test responses, including refusals, and excluded queue waits, retry backoff and failed attempts. These measurements don’t isolate inference time, and the separately collected historical timings don’t establish a controlled speed advantage.
This is an initial transferred-prompt comparison on benchmarks whose earlier results were already known. It supports retaining Jev as the stronger measured baseline in this setup. Comparable tuning budgets and fresh held-out material would be needed to compare the systems’ attainable performance.
What this means in practice
Tuning wasn’t equal: Jev had prompt and threshold selection, while Luna’s later Chat search covered three prompts on ToxicChat only. Decisions received mechanically transferred prompts without its own search or fitted cutoff. I knew the earlier test results when I designed the follow-ups, and benchmark exposure during model training is unknown. These exploratory comparisons aren’t an equally tuned model leaderboard.
A good F1 also doesn’t guarantee safe automatic approval. A separate post-test Aegis what-if reused cached unsafe scores, allowing messages at p ≤ 0.05, blocking them at p ≥ 0.95 and deferring the rest. With the baseline prompt, 52 of 560 automatically allowed messages had unsafe labels (9.29%), while 969 of 1,923 messages remained deferred and unresolved. These were illustrative thresholds, not a selected deployment policy. No human review or multi-stage workflow was tested, and the scores weren’t validated probabilities of correctness.
Jev remains a promising candidate for this task, and the initial Decisions comparison doesn’t give me a reason to replace the existing baseline. Before deployment, I would define the application’s safety policy and refusal fallback, test fresh messages and examine missed harms and false alarms separately.
Methods and sources
The methods, prompts, aggregate results and analysis code for the original comparisons are available in the experiment repository, along with verification code. The Decisions follow-up uses the completed 7 October experiment and its offline refusal and latency checks. OpenAI documents the endpoint in its Decisions guide and API reference.