Research
Updated · 31 min read
AIcontent moderationresearch

Evaluating Jev for content moderation

A comparison of Jev, GPT-5.6 Luna and GPT-6 Luna on four moderation datasets, including prompt trade-offs, historical cost and recorded request time.

Forest-green ink and ochre flecks flow through three fine, curved mesh filters on textured ivory paper.

A moderation check has two jobs: catch harmful requests and let harmless ones through. Blocking an ordinary question about a sensitive subject is a mistake too. Can Jev handle that decision well enough to be useful, and do it faster and cheaper?

The original tests compared Jev 1.13.0, GPT-5.6 Luna and GPT-6 Luna on four public datasets. Jev using the baseline prompt had the highest observed F1 on all four, although some differences were too uncertain to call. A later initial comparison with GPT-6 Luna through OpenAI Decisions also favoured Jev in this setup, with important differences in tuning and refusal handling.

What the datasets test

  • ToxicChat contains real chatbot requests labelled for toxicity, testing the messy language people actually use.

  • XSTest pairs harmless prompts that can look unsafe with genuinely unsafe ones, testing whether a model overreacts to wording.

  • Aegis 2 is NVIDIA’s content-safety dataset, covering a broad range of safety risks under a detailed labelling policy.

  • WildGuardTest includes ordinary and deliberately adversarial requests, testing harmful intent as well as misleading framing.

These aren’t interchangeable definitions of safety. An abusive message and a politely worded request for harmful help pose different problems.

How the models were tested

Each model classified message text as safe or unsafe, checked against dataset labels. The original tests didn’t assess generated answers or refusal behaviour. After exclusions, each Luna comparison scored 9,155 rows from 9,035 distinct API payloads. Exact duplicate payloads reused predictions while every eligible row kept its scoring weight. The later Decisions comparison and its classification refusals are discussed separately below.

Selected development and test inputs were separated by Unicode, case and whitespace-normalised exact text. A later similarity audit found 18 ToxicChat test rows, eight toxic, near development inputs. Excluding them in a post-test sensitivity check preserved the main conclusions. This doesn’t establish full semantic independence.

Jev returned a safety score, converted to a decision using a 0.54 cutoff. Here, Jev baseline prompt and Jev content-focused prompt distinguish two prompts for the same Jev 1.13.0 model, with the same cutoff.

The baseline prompt focuses on whether the message asks for harmful assistance. The content-focused prompt additionally considers harmful content in the message itself. This broadens the decision beyond the help requested, so the choice of prompt can change what the moderation check catches.

I chose the baseline prompt and cutoff using 200 WildGuard training examples, then checked them on an independent 200-example sample before the held-out tests. Both Luna models used the same safety criteria in a strict Boolean JSON prompt, returning true or false with reasoning disabled.

Both Jev prompts were frozen before the 1,923-row Aegis test and used the 0.54 cutoff without further tuning.

A false positive is a harmless message wrongly flagged. A false negative is a harmful message missed. F1 balances how often flagged messages really are unsafe with how many unsafe messages the model catches. It isn’t overall accuracy.

What the results show

DatasetTest rowsJev baseline promptGPT-5.6 LunaGPT-6 Luna
ToxicChat5,08375.4066.8663.68
XSTest45093.4088.2791.46
Aegis 21,92380.2779.1277.58
WildGuardTest1,69988.3083.5485.17

F1 for unsafe or toxic messages, on a 0–100 scale. Higher is better. The Jev column uses the baseline prompt throughout.

Sensitivity checks used 2,000 paired bootstrap resamples of normalised-message groups. The Aegis comparison with GPT-5.6 and XSTest with GPT-6 were inconclusive. The other paired differences favoured Jev. These nominal intervals describe sampling uncertainty conditional on cached predictions. They weren’t adjusted for multiple comparisons and don’t capture prompt selection, repeated model runs or all shared-template dependence.

“Jev using the baseline prompt had the highest observed F1 on all four, although some differences were too uncertain to call.”

F1 across four moderation benchmarks ToxicChat: Jev 75.40%, GPT-5.6 Luna 66.86%, GPT-6 Luna 63.68%. XSTest: Jev 93.40%, GPT-5.6 Luna 88.27%, GPT-6 Luna 91.46%. Aegis 2: Jev 80.27%, GPT-5.6 Luna 79.12%, GPT-6 Luna 77.58%. WildGuardTest: Jev 88.30%, GPT-5.6 Luna 83.54%, GPT-6 Luna 85.17%. Observed F1 point estimates. Aegis versus GPT-5.6 and XSTest versus GPT-6 were inconclusive in the sampling checks. Figure 1 F1 across four moderation benchmarks F1 on labelled test rows · higher is better Jev uses the original policy ToxicChat 5,083 rows 0% 50% 100% Jev 75.40% GPT-5.6 Luna 66.86% GPT-6 Luna 63.68% XSTest 450 rows 0% 50% 100% Jev 93.40% GPT-5.6 Luna 88.27% GPT-6 Luna 91.46% Aegis 2 1,923 rows 0% 50% 100% Jev 80.27% GPT-5.6 Luna 79.12% GPT-6 Luna 77.58% WildGuardTest 1,699 rows 0% 50% 100% Jev 88.30% GPT-5.6 Luna 83.54% GPT-6 Luna 85.17% F1 (%) F1 across four moderation benchmarks ToxicChat: Jev 75.40%, GPT-5.6 Luna 66.86%, GPT-6 Luna 63.68%. XSTest: Jev 93.40%, GPT-5.6 Luna 88.27%, GPT-6 Luna 91.46%. Aegis 2: Jev 80.27%, GPT-5.6 Luna 79.12%, GPT-6 Luna 77.58%. WildGuardTest: Jev 88.30%, GPT-5.6 Luna 83.54%, GPT-6 Luna 85.17%. Observed F1 point estimates. Aegis versus GPT-5.6 and XSTest versus GPT-6 were inconclusive in the sampling checks. Figure 1 F1 across four moderation benchmarks F1 on labelled test rows · higher is better Jev uses the original policy ToxicChat 5,083 rows 0% 50% 100% Jev 75.40% GPT-5.6 Luna 66.86% GPT-6 Luna 63.68% XSTest 450 rows 0% 50% 100% Jev 93.40% GPT-5.6 Luna 88.27% GPT-6 Luna 91.46% Aegis 2 1,923 rows 0% 50% 100% Jev 80.27% GPT-5.6 Luna 79.12% GPT-6 Luna 77.58% WildGuardTest 1,699 rows 0% 50% 100% Jev 88.30% GPT-5.6 Luna 83.54% GPT-6 Luna 85.17% F1 (%)

Figure 1. Baseline benchmark results, unchanged after the later GPT-6 prompt search. The charts’ original Jev results refer to the Jev baseline prompt.

On WildGuard, Jev caught 638 of 754 harmful requests, versus GPT-6’s 600, and falsely flagged 53 harmless messages versus 55. That’s 38 more harmful requests caught with two fewer false alarms. Jev still missed 116. GPT-6 improved on GPT-5.6 for XSTest and WildGuard, but fell behind on ToxicChat and Aegis.

What prompt changes showed

I selected the Jev content-focused prompt from three prompts tested on 1,000 ToxicChat training groups, then checked it on another 1,000 independent groups. The selected prompt was fixed before test inference and the 0.54 cutoff stayed unchanged.

DatasetJev baseline promptJev content-focused prompt
ToxicChat75.4080.23
XSTest93.4090.72
Aegis 280.2780.77
WildGuardTest88.30Not run

F1 for the same Jev 1.13.0 model with two prompts, on a 0–100 scale. Higher is better.

The ToxicChat gain and XSTest decline were both supported by the uncertainty checks. The smaller Aegis change was inconclusive, and the content-focused prompt wasn’t run on WildGuardTest. This suggests that improving the fit to one safety definition can worsen the fit to another.

I kept the baseline in the main comparison because it has results on all four datasets and uses the same safety criteria carried into both Luna prompts. Choosing whichever Jev prompt scored best on each test set would be a different comparison. The content-focused prompt shows a trade-off, rather than a consistently better replacement.

“The ToxicChat gain and XSTest decline were both supported by the uncertainty checks.”

The GPT-6 Luna search tested three fixed prompts on 600 previously unused ToxicChat training groups. The original won with 69.70 F1, versus 68.75 and 66.67. On 600 separate confirmation groups, it scored 72.16 F1, wrongly flagging 2% of harmless messages. Following the pre-agreed protocol, I reused its historical four-dataset predictions: there was no full benchmark rerun. This small search found no improvement, rather than proving none is possible.

Cost and recorded request time

Historical estimated cost per thousand requests was $0.019–$0.022 for Jev, $0.060–$0.077 for GPT-5.6 and $0.029–$0.037 for GPT-6. These estimates use returned token usage and recorded rates, including cache writes.

Estimated cost across the completed runs ToxicChat: Jev $0.0201, GPT-5.6 Luna $0.0661, GPT-6 Luna $0.0318. XSTest: Jev $0.0189, GPT-5.6 Luna $0.0604, GPT-6 Luna $0.0289. Aegis 2: Jev $0.0210, GPT-5.6 Luna $0.0708, GPT-6 Luna $0.0345. WildGuardTest: Jev $0.0224, GPT-5.6 Luna $0.0766, GPT-6 Luna $0.0370. Historical estimates. Aegis covers 1,915 Jev requests and 1,910 requests for each Luna run. Figure 2 Estimated cost across the completed runs Historical USD per 1,000 requests Jev uses the original policy ToxicChat $0.00 $0.04 $0.08 Jev $0.0201 GPT-5.6 Luna $0.0661 GPT-6 Luna $0.0318 XSTest $0.00 $0.04 $0.08 Jev $0.0189 GPT-5.6 Luna $0.0604 GPT-6 Luna $0.0289 Aegis 2 $0.00 $0.04 $0.08 Jev $0.0210 GPT-5.6 Luna $0.0708 GPT-6 Luna $0.0345 WildGuardTest $0.00 $0.04 $0.08 Jev $0.0224 GPT-5.6 Luna $0.0766 GPT-6 Luna $0.0370 Estimated USD per 1,000 requests Estimated cost across the completed runs ToxicChat: Jev $0.0201, GPT-5.6 Luna $0.0661, GPT-6 Luna $0.0318. XSTest: Jev $0.0189, GPT-5.6 Luna $0.0604, GPT-6 Luna $0.0289. Aegis 2: Jev $0.0210, GPT-5.6 Luna $0.0708, GPT-6 Luna $0.0345. WildGuardTest: Jev $0.0224, GPT-5.6 Luna $0.0766, GPT-6 Luna $0.0370. Historical estimates. Aegis covers 1,915 Jev requests and 1,910 requests for each Luna run. Figure 2 Estimated cost across the completed runs Historical USD per 1,000 requests Jev uses the original policy ToxicChat $0.00 $0.04 $0.08 Jev $0.0201 GPT-5.6 Luna $0.0661 GPT-6 Luna $0.0318 XSTest $0.00 $0.04 $0.08 Jev $0.0189 GPT-5.6 Luna $0.0604 GPT-6 Luna $0.0289 Aegis 2 $0.00 $0.04 $0.08 Jev $0.0210 GPT-5.6 Luna $0.0708 GPT-6 Luna $0.0345 WildGuardTest $0.00 $0.04 $0.08 Jev $0.0224 GPT-5.6 Luna $0.0766 GPT-6 Luna $0.0370 Estimated USD per 1,000 requests

Figure 2. Cost per 1,000 unique requests. Counts: ToxicChat 4,976, XSTest 450, WildGuard 1,699, and Aegis 1,915 for Jev and 1,910 for each Luna run.

Median recorded times for successful requests were 289–316 milliseconds for Jev, 707–720 for GPT-5.6 and 750–770 for GPT-6. Jev’s timer stopped when the HTTP call returned. Luna’s also included JSON parsing, response-file writing and label/cost processing. The models ran separately under different service conditions, and the saved timings can’t isolate local processing costs. This wasn’t a controlled inference-speed comparison.

Recorded time for successful requests ToxicChat: Jev 305, GPT-5.6 Luna 715, GPT-6 Luna 754. XSTest: Jev 289, GPT-5.6 Luna 707, GPT-6 Luna 750. Aegis 2: Jev 290, GPT-5.6 Luna 710, GPT-6 Luna 760. WildGuardTest: Jev 316, GPT-5.6 Luna 720, GPT-6 Luna 770. Recorded successful-request times from separate runs with different network and service conditions. Jev stops at HTTP return. Luna includes JSON parsing, response-file writing and label/cost processing. Client queue and rate-limit waiting are excluded. No timing uncertainty intervals were saved. Figure 3 Recorded time for successful requests Successful requests · milliseconds Jev uses the original policy ToxicChat 0 400 800 Jev 305 GPT-5.6 Luna 715 GPT-6 Luna 754 XSTest 0 400 800 Jev 289 GPT-5.6 Luna 707 GPT-6 Luna 750 Aegis 2 0 400 800 Jev 290 GPT-5.6 Luna 710 GPT-6 Luna 760 WildGuardTest 0 400 800 Jev 316 GPT-5.6 Luna 720 GPT-6 Luna 770 Median recorded request time (ms) Recorded time for successful requests ToxicChat: Jev 305, GPT-5.6 Luna 715, GPT-6 Luna 754. XSTest: Jev 289, GPT-5.6 Luna 707, GPT-6 Luna 750. Aegis 2: Jev 290, GPT-5.6 Luna 710, GPT-6 Luna 760. WildGuardTest: Jev 316, GPT-5.6 Luna 720, GPT-6 Luna 770. Recorded successful-request times from separate runs with different network and service conditions. Jev stops at HTTP return. Luna includes JSON parsing, response-file writing and label/cost processing. Client queue and rate-limit waiting are excluded. No timing uncertainty intervals were saved. Figure 3 Recorded time for successful requests Successful requests · milliseconds Jev uses the original policy ToxicChat 0 400 800 Jev 305 GPT-5.6 Luna 715 GPT-6 Luna 754 XSTest 0 400 800 Jev 289 GPT-5.6 Luna 707 GPT-6 Luna 750 Aegis 2 0 400 800 Jev 290 GPT-5.6 Luna 710 GPT-6 Luna 760 WildGuardTest 0 400 800 Jev 316 GPT-5.6 Luna 720 GPT-6 Luna 770 Median recorded request time (ms)

Figure 3. Median recorded time for successful requests. Client queue and rate-limit waiting are excluded. No timing uncertainty intervals were saved.

An initial comparison with OpenAI Decisions

I also tested GPT-6 Luna through OpenAI’s Decisions endpoint. Two frozen prompts transferred the existing criteria: the original assistant-safety prompt, with its Chat output-format instruction removed, and a policy variant formatted from the Jev content-focused instructions. Each scored the same 9,155 rows from 9,035 distinct requests, using a 0.5 probability cutoff. Historical Jev and Chat predictions were reused. No Decisions prompt search or threshold fitting was performed, and no further tuning was conducted.

DatasetJev baseline promptDecisions originalDecisions policy variant
ToxicChat75.4050.0062.88
XSTest93.4086.5084.32
Aegis 280.2772.9972.41
WildGuardTest88.3079.6373.08

F1 on a 0–100 scale. Decisions scores include blocking requests when the endpoint refused to classify them.

Decisions, Jev and baseline Luna ToxicChat, 5,083 scored rows: Jev · baseline 75.40, nominal 95% interval 71.65 to 78.99, GPT-6 Luna · Chat 63.68, nominal 95% interval 59.15 to 67.84, Decisions · original 50.00, nominal 95% interval 44.44 to 55.39, Decisions · policy variant 62.88, nominal 95% interval 58.12 to 67.25. Decisions probability coverage: original 100.00%, 5,083 of 5,083 rows, policy 99.74%, 5,070 of 5,083 rows. XSTest, 450 scored rows: Jev · baseline 93.40, nominal 95% interval 90.77 to 95.67, GPT-6 Luna · Chat 91.46, nominal 95% interval 88.40 to 94.38, Decisions · original 86.50, nominal 95% interval 82.62 to 90.13, Decisions · policy variant 84.32, nominal 95% interval 79.89 to 88.22. Decisions probability coverage: original 98.67%, 444 of 450 rows, policy 90.44%, 407 of 450 rows. Aegis 2, 1,923 scored rows: Jev · baseline 80.27, nominal 95% interval 78.27 to 82.10, GPT-6 Luna · Chat 77.58, nominal 95% interval 75.42 to 79.66, Decisions · original 72.99, nominal 95% interval 70.48 to 75.28, Decisions · policy variant 72.41, nominal 95% interval 70.01 to 74.88. Decisions probability coverage: original 98.75%, 1,899 of 1,923 rows, policy 95.06%, 1,828 of 1,923 rows. WildGuardTest, 1,699 scored rows: Jev · baseline 88.30, nominal 95% interval 86.46 to 89.97, GPT-6 Luna · Chat 85.17, nominal 95% interval 83.12 to 87.12, Decisions · original 79.63, nominal 95% interval 77.16 to 81.95, Decisions · policy variant 73.08, nominal 95% interval 70.03 to 75.69. Decisions probability coverage: original 99.82%, 1,696 of 1,699 rows, policy 97.41%, 1,655 of 1,699 rows. Figure 4. Unsafe-positive F1 on the same full scored rows, with Decisions refusals blocked. GPT-6 Luna Chat uses the unchanged Boolean baseline prompt. Jev uses its selected baseline prompt and 0.54 cutoff. Decisions uses the original and policy-variant prompts at 0.5. Whiskers show recorded nominal 95% grouped-bootstrap intervals for the saved predictions, without correction for multiple comparisons. These are not equally tuned systems, and interval overlap is not a paired significance test. Figure 4 Decisions, Jev and baseline Luna Full-row F1 · higher is better · Decisions refusals counted as blocks Recorded nominal 95% sampling intervals · unequal tuning ToxicChat 5,083 rows Jev · baseline 75.40 GPT-6 Luna · Chat 63.68 Decisions · original 50.00 Decisions · policy variant 62.88 0 50 100 Decisions probability coverage: original 100.00% · policy variant 99.74% XSTest 450 rows Jev · baseline 93.40 GPT-6 Luna · Chat 91.46 Decisions · original 86.50 Decisions · policy variant 84.32 0 50 100 Decisions probability coverage: original 98.67% · policy variant 90.44% Aegis 2 1,923 rows Jev · baseline 80.27 GPT-6 Luna · Chat 77.58 Decisions · original 72.99 Decisions · policy variant 72.41 0 50 100 Decisions probability coverage: original 98.75% · policy variant 95.06% WildGuardTest 1,699 rows Jev · baseline 88.30 GPT-6 Luna · Chat 85.17 Decisions · original 79.63 Decisions · policy variant 73.08 0 50 100 Decisions probability coverage: original 99.82% · policy variant 97.41% Unsafe-positive F1 (%) Decisions, Jev and baseline Luna ToxicChat, 5,083 scored rows: Jev · baseline 75.40, nominal 95% interval 71.65 to 78.99, GPT-6 Luna · Chat 63.68, nominal 95% interval 59.15 to 67.84, Decisions · original 50.00, nominal 95% interval 44.44 to 55.39, Decisions · policy variant 62.88, nominal 95% interval 58.12 to 67.25. Decisions probability coverage: original 100.00%, 5,083 of 5,083 rows, policy 99.74%, 5,070 of 5,083 rows. XSTest, 450 scored rows: Jev · baseline 93.40, nominal 95% interval 90.77 to 95.67, GPT-6 Luna · Chat 91.46, nominal 95% interval 88.40 to 94.38, Decisions · original 86.50, nominal 95% interval 82.62 to 90.13, Decisions · policy variant 84.32, nominal 95% interval 79.89 to 88.22. Decisions probability coverage: original 98.67%, 444 of 450 rows, policy 90.44%, 407 of 450 rows. Aegis 2, 1,923 scored rows: Jev · baseline 80.27, nominal 95% interval 78.27 to 82.10, GPT-6 Luna · Chat 77.58, nominal 95% interval 75.42 to 79.66, Decisions · original 72.99, nominal 95% interval 70.48 to 75.28, Decisions · policy variant 72.41, nominal 95% interval 70.01 to 74.88. Decisions probability coverage: original 98.75%, 1,899 of 1,923 rows, policy 95.06%, 1,828 of 1,923 rows. WildGuardTest, 1,699 scored rows: Jev · baseline 88.30, nominal 95% interval 86.46 to 89.97, GPT-6 Luna · Chat 85.17, nominal 95% interval 83.12 to 87.12, Decisions · original 79.63, nominal 95% interval 77.16 to 81.95, Decisions · policy variant 73.08, nominal 95% interval 70.03 to 75.69. Decisions probability coverage: original 99.82%, 1,696 of 1,699 rows, policy 97.41%, 1,655 of 1,699 rows. Figure 4. Unsafe-positive F1 on the same full scored rows, with Decisions refusals blocked. GPT-6 Luna Chat uses the unchanged Boolean baseline prompt. Jev uses its selected baseline prompt and 0.54 cutoff. Decisions uses the original and policy-variant prompts at 0.5. Whiskers show recorded nominal 95% grouped-bootstrap intervals for the saved predictions, without correction for multiple comparisons. These are not equally tuned systems, and interval overlap is not a paired significance test. Figure 4 Decisions, Jev and baseline Luna Full-row F1 · higher is better Decisions refusals counted as blocks Nominal 95% intervals · unequal tuning ToxicChat 5,083 rows Jev · baseline 75.40 GPT-6 Luna · Chat 63.68 Decisions · original 50.00 Decisions · policy variant 62.88 0 50 100 Probability coverage · original 100.00% policy variant 99.74% XSTest 450 rows Jev · baseline 93.40 GPT-6 Luna · Chat 91.46 Decisions · original 86.50 Decisions · policy variant 84.32 0 50 100 Probability coverage · original 98.67% policy variant 90.44% Aegis 2 1,923 rows Jev · baseline 80.27 GPT-6 Luna · Chat 77.58 Decisions · original 72.99 Decisions · policy variant 72.41 0 50 100 Probability coverage · original 98.75% policy variant 95.06% WildGuardTest 1,699 rows Jev · baseline 88.30 GPT-6 Luna · Chat 85.17 Decisions · original 79.63 Decisions · policy variant 73.08 0 50 100 Probability coverage · original 99.82% policy variant 97.41% Unsafe-positive F1 (%)
Figure 4. Unsafe-positive F1 on the same full scored rows, with Decisions refusals blocked. GPT-6 Luna Chat uses the unchanged Boolean baseline prompt. Jev uses its selected baseline prompt and 0.54 cutoff. Decisions uses the original and policy-variant prompts at 0.5. Whiskers show recorded nominal 95% grouped-bootstrap intervals for the saved predictions, without correction for multiple comparisons. These are not equally tuned systems, and interval overlap is not a paired significance test.

Jev had higher observed F1 in all eight comparisons. The nominal 95% intervals from 2,000 paired resamples of normalised-message groups favoured Jev throughout, with the same sampling and multiple-comparison limitations described above. The policy variant improved Decisions on ToxicChat but reduced F1 elsewhere. Original Decisions caught 36.74% of toxic ToxicChat messages, compared with Jev’s 71.55%, while falsely flagging 0.78% of harmless messages, compared with 1.40%. Lower false-positive rates came with more missed harms.

Refusals need separate attention because they contain no classification probability. The primary analysis counted them as blocked deployment actions, so it measures the classifier plus that fallback. This departed from the earlier preflight plan to leave refusals unresolved, although the blocking rule was frozen before inference. Probability coverage was 99.64% of scored rows for the original prompt and 97.87% for the policy variant. On XSTest, the variant refused 9.56% of rows: allowing rather than blocking those requests reduced F1 from 84.32 to 72.78. An offline check excluded refusals and compared each configuration on exactly the same answered rows. Jev remained ahead in all eight comparisons, but those selected subsets differ between prompts and no new confidence intervals were calculated.

The run’s measured token usage gave a $0.6790464 base cost estimate, including the development pilot. A conservative ledger allowance totalled $1.0197696, including a 50% margin and a reservation for one failed attempt. Neither figure is an invoice-reconciled bill. Base estimates across datasets and prompts were $0.032–$0.043 per thousand unique test requests.

Median client request times were around 301 milliseconds, including network, server and local response processing. They used final valid test responses, including refusals, and excluded queue waits, retry backoff and failed attempts. These measurements don’t isolate inference time, and the separately collected historical timings don’t establish a controlled speed advantage.

This is an initial transferred-prompt comparison on benchmarks whose earlier results were already known. It supports retaining Jev as the stronger measured baseline in this setup. Comparable tuning budgets and fresh held-out material would be needed to compare the systems’ attainable performance.

What this means in practice

Tuning wasn’t equal: Jev had prompt and threshold selection, while Luna’s later Chat search covered three prompts on ToxicChat only. Decisions received mechanically transferred prompts without its own search or fitted cutoff. I knew the earlier test results when I designed the follow-ups, and benchmark exposure during model training is unknown. These exploratory comparisons aren’t an equally tuned model leaderboard.

A good F1 also doesn’t guarantee safe automatic approval. A separate post-test Aegis what-if reused cached unsafe scores, allowing messages at p ≤ 0.05, blocking them at p ≥ 0.95 and deferring the rest. With the baseline prompt, 52 of 560 automatically allowed messages had unsafe labels (9.29%), while 969 of 1,923 messages remained deferred and unresolved. These were illustrative thresholds, not a selected deployment policy. No human review or multi-stage workflow was tested, and the scores weren’t validated probabilities of correctness.

Jev remains a promising candidate for this task, and the initial Decisions comparison doesn’t give me a reason to replace the existing baseline. Before deployment, I would define the application’s safety policy and refusal fallback, test fresh messages and examine missed harms and false alarms separately.

Methods and sources

The methods, prompts, aggregate results and analysis code for the original comparisons are available in the experiment repository, along with verification code. The Decisions follow-up uses the completed 7 October experiment and its offline refusal and latency checks. OpenAI documents the endpoint in its Decisions guide and API reference.