Part 2/3 Owning Intelligence Series: How Coinbase Built, Evaluated, and Post-Trained a Fraud Agent

โดย Yao Ma, Xuwei Tan

TL;DR: Newer LLMs are not necessarily better at domain-specific tasks. On our fraud benchmark, newer Opus, Sonnet, and GPT versions all had lower recall, F1, and dollar-weighted recall—even when precision improved. The lesson: choose models using high-quality, domain-specific benchmarks, with reproducible evaluation and ideally deterministic scoring against verified outcomes. These findings also motivated our next step: post-training open-weight models for greater control over quality, cost, and deployment.

Coinbase Logo

Part 2 — Evaluate: Why Newer Models Aren’t Always Better for Your Domain

In This Series

  • Part 1 — Build: We layer an LLM risk agent onto Onramp’s existing ML model and rules. Online A/B experiment shows 30% fewer fraudulent transactions and 22% less fraud value.

  • Part 2 — Evaluate: We compared Opus, Sonnet, and GPT releases on our fraud benchmark. Each newer version scored lower on recall and F1, reinforcing the need for domain-specific evaluation.

  • Part 3 — Own: Using SkyRL on Anyscale, we post-trained Qwen3.5-9B on our proprietary Onramp fraud dataset. The resulting model outperformed Opus 4.5 on all four fraud-detection metrics and delivered 55% lower end-to-end online serving latency.

The Challenge: A Model Upgrade Can Change the Risk Policy

In Part 1, we described an LLM-powered agent that reviews recent transaction behavior after the traditional model and rules have completed their checks. An online experiment showed that this additional review could reduce fraud without replacing the existing risk stack.

That left a separate question: which model should power the agent, and when should we upgrade it?

A model change can look like a small configuration update. The application still assembles transaction context, asks for a risk classification, and applies the same decision policy. Yet the meaning of that classification can shift. A newer model may interpret the same evidence more conservatively, react more strongly to a particular pattern, or draw a different boundary between suspicious and legitimate activity.

For a payment-risk system, that is a behavioral change—not just an infrastructure update. General-purpose benchmarks do not tell us how it will affect missed fraud, unnecessary blocks, or the value of fraud detected. We needed to measure those outcomes on our own task.

The Benchmark: A Consistent Test Across Models and Versions

We evaluated a fixed benchmark of 16,140 transactions across 7,293 users, drawn from nine weeks of production transaction data before the risk agent described in Part 1 was rolled out. We chose this pre-rollout period to avoid bias from the agent’s own interventions: its blocking decisions could otherwise change which transactions proceeded and which fraud outcomes became observable. The cohort contains 813 confirmed fraudulent transactions, retaining all matured fraud cases from the source window while sampling legitimate traffic.

Unlike Part 1’s online experiment, this evaluation replays historical transactions with known outcomes. Its purpose is to compare candidate decision models under a common evaluation setup, not to measure the live business impact of deploying each one.

1. Keep the Evaluation Task Consistent

The version comparison focuses on the decision agent alone: the model evaluates a transaction using its recent behavioral context and returns a risk classification. Its decision guidance remains fixed during evaluation rather than adapting to newly observed outcomes.

Each candidate is evaluated on the same transaction cohort using the contextual-review procedure and a fixed risk-to-decision policy. This asks a practical upgrade question: how does a different model behave in the existing decision setup?

2. Measure Both Detection Quality and Financial Coverage

We compare four complementary metrics:

  • Precision: among transactions classified as fraud, how many were actually fraudulent?

  • Recall: among fraudulent transactions, how many did the agent identify?

  • F1: the harmonic mean of precision and recall, summarizing their balance.

  • Dollar-weighted recall: what share of the total fraudulent transaction value did the agent identify?

The last metric is particularly important for payment risk. Catching more incidents and catching more fraud value are related, but they are not the same objective. A model can miss relatively few transactions yet fail to detect a disproportionate share of the financial exposure.

Results: Newer Model Versions Did Not Improve Fraud Detection

We compare Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Across all three pairs, the newer version had lower recall, F1, and dollar-weighted recall.

evaluate 1

Figure 1: Change from the earlier to the newer version within each model family. Values are percentage-point differences measured on the same benchmark.

Opus and Sonnet: Regression Across All Four Metrics

Opus 5 and Sonnet 5 scored lower than their respective earlier versions on every reported metric. The size of the change differed substantially: Opus recall fell by 0.8 percentage points, while Sonnet recall fell by 22.2 points. Sonnet’s dollar-weighted recall also fell by 22.9 points.

This was not simply a trade-off in which better precision compensated for lower recall. Both precision and recall declined. Under the evaluated decision setup, the newer versions were worse at distinguishing fraud from legitimate activity.

GPT: Higher Precision, Lower Fraud Coverage

GPT-5.6 (sol) showed a different pattern. Precision improved by 11.5 percentage points, but recall fell by 20.7 points and dollar-weighted recall fell by 21.8 points. F1 also declined.

Why Can a Newer Model Perform Worse?

Our benchmark shows a regression, not its cause. Without access to providers’ training data or objectives, we can identify plausible explanations but cannot attribute the changes to a specific training decision:

  • Decision boundaries and prompt behavior can shift. The same risk label may mean something different across versions. Our GPT results—higher precision but lower recall—are consistent with more conservative classification. Chen et al. (2023) also documented changes in instruction following and output formatting across model snapshots.

  • Post-training can introduce trade-offs. Ouyang et al. (2022) found that preference optimization reduced performance on some benchmarks, although changes to the training objective mitigated those regressions. Kotha et al. (2023) showed that some apparent capability losses after fine-tuning could be recovered through prompting—suggesting changed behavior rather than erased knowledge.

  • Data mixtures and task priorities can change. DoReMi (Xie et al., 2023) demonstrates that training-data composition affects downstream performance.

What This Means for Model Selection

The results change how we frame an upgrade. Instead of asking which release is newest, the useful question is whether a candidate improves the outcomes that matter under the constraints of the agentic system.

A practical evaluation process should answer three questions:

  • Does the candidate improve the overall trade-off? Review precision, recall, and dollar-weighted recall together. A gain in one metric should not conceal a larger loss elsewhere.

  • Is the comparison between models or between policies? First test the candidate under the existing decision setup. If it needs new prompts or thresholds, evaluate that adapted configuration separately rather than attributing the result to the model alone.

  • Can it meet the serving requirements? Detection quality is only one part of suitability. Latency, reliability, and operating cost still determine whether the model fits the payment flow.

From Owning the Evaluation to Owning the Model

Closed-source models helped us build and deploy a useful fraud-review agent. This benchmark study did not invalidate that approach. It showed why we needed our own evidence to decide which model to use—and why improvements in general capability could not substitute for improvements on our task.

It also exposed a limit. Evaluation lets us detect a regression, retain a stronger candidate, or adapt the surrounding system. It does not let us change the provider’s weights or training process.

Those results motivated our move toward open-weight models. The question became whether domain-specific post-training could give us the fraud-detection behavior we wanted, with greater control over the model and its serving environment.

Owning the evaluation gave us a way to detect regressions and choose models based on evidence. The next step was to gain more control over the model itself. In Part 3, we use SkyRL on Anyscale to post-train Qwen3.5-9B and evaluate its detection quality and serving latency against Opus 4.5.


เรื่องล่าสุด

ข้อจำกัดความรับผิดชอบ: การเทรดอนุพันธ์ผ่านแพลตฟอร์ม Coinbase Advanced นำเสนอแก่ลูกค้า EEA ที่มีสิทธิ์ โดย Coinbase Financial Services Europe Ltd. (ใบอนุญาต CySEC 374/19) ในการเข้าถึงอนุพันธ์ ลูกค้าจะต้องผ่านการตรวจสอบมาตรฐานของเราเพื่อพิจารณาคุณสมบัติและความเหมาะสมสำหรับผลิตภัณฑ์นี้