Part 3/3 Owning Intelligence Series: How Coinbase Built, Evaluated, and Post-Trained a Fraud Agent

di Yao Ma, Xuwei Tan, Akshit Trehan, Aman Choudhary, Anyscale Team: Eric Tang, Sumanth Hegde, Seiji Eicher, Jess Kong

TL;DR: Using SkyRL on Anyscale, we post-trained Qwen3.5-9B for Onramp fraud detection and served it with Ray Serve LLM and vLLM. On our benchmark, it exceeded Opus 4.5 across all four detection metrics, including 9.6% higher F1. In online inference, we reduced end to end latency by 55%.

Coinbase Logo

In This Series

  • Part 1 — Build: We layer an LLM risk agent onto Onramp’s existing ML model and rules. Online A/B experiment shows 30% fewer fraudulent transactions and 22% less fraud value.

  • Part 2 — Evaluate: We compared Opus, Sonnet, and GPT releases on our fraud benchmark. Each newer version scored lower on recall and F1, reinforcing the need for domain-specific evaluation.

  • Part 3 — Own: Using SkyRL on Anyscale, we post-trained Qwen3.5-9B on our proprietary Onramp fraud dataset. The resulting model outperformed Opus 4.5 on all four fraud-detection metrics and delivered 55% lower end-to-end online serving latency.

From Evaluating Models to Improving Them

Our domain-specific benchmark helped identify regressions, but with Opus 4.5 our control stopped at the API boundary. We could refine prompts, improve the agent harness, or choose another model version—but we could not adapt the model’s weights to Onramp’s fraud patterns.

That raised a practical question: could we specialize and self-host an open-weight model without sacrificing detection quality? We focused on one narrow task: classifying transaction risk from recent behavioral context.

The approach brought together Coinbase’s proprietary fraud data and outcome-based rewards, SkyRL’s reinforcement-learning framework, and Anyscale’s platform for running the training workload—giving us a path from evaluating model behavior to improving it.

own 1

Figure 1: The Onramp fraud agent’s training-to-serving workflow on Anyscale. SkyRL post-trains Qwen3.5-9B using Coinbase’s data and outcome-based rewards. Ray Serve LLM and vLLM serve the merged checkpoint, while the agent harness translates risk classifications into decisions.

Post-Training for the Decision We Actually Need

We started with Qwen3.5-9B, an open-weight model with nine billion parameters. Its size made single-GPU serving feasible. 

The decision task remained the same as in Part 2: review transaction context and return a risk classification. The harness layer—not the model—remained responsible for translating that classification into a decision. To minimize online serving latency, we separate the training-time response format from the serving contract: the payment flow needs only a compact risk verdict, not a written explanation.

Learning from Fraud Outcomes with RLVR

We used Reinforcement Learning with Verifiable Rewards (RLVR): scoring the model’s risk classifications against known historical fraud outcomes using a deterministic reward function. Training prompts followed the same transaction format used at evaluation. 

We used Group Relative Policy Optimization (GRPO), a reinforcement-learning method that compares multiple sampled responses to each prompt and updates the model toward responses with higher rewards. A custom fraud environment in SkyRL parses each response and computes an outcome-based reward, with additional checks for response formatting. The scoring is deterministic: it does not depend on another LLM judging how persuasive the response sounds.

Fraud detection also requires care with class imbalance. A reward that favors the majority class can produce a detector that appears accurate while missing the events that matter. Our reward balanced the contributions of fraudulent and legitimate examples rather than rewarding a constant prediction strategy.

We applied these updates using Low-Rank Adaptation (LoRA), which trains a compact set of additional parameters instead of updating every model weight. This gave us a practical way to specialize the model while retaining the pretrained backbone.

Run the Training Loop with SkyRL on Anyscale

SkyRL connects response generation, reward computation, and model updates. Our training configuration uses its Megatron training backend with LoRA, vLLM for generating candidate rollouts, and a custom fraud-scoring environment. This lets us specialize the reward and data for our task without building the reinforcement-learning loop from scratch. With SkyRL on Anyscale, we were able to train the Qwen3.5-9B model on just one NVIDIA RTX PRO 6000 96GB GPU for a total cost of under $100.

Results: The Improvement Came from Domain Adaptation

We compared post-trained Qwen3.5-9B with its unadapted base and Opus 4.5 on our fraud benchmark. Each operated as a single-pass decision agent, without adaptive rules or reflection, to isolate the model’s contribution.

The unadapted Qwen model performed poorly; post-training improved the same backbone across all four detection metrics. Opus 4.5 also provides the baseline for the serving comparison below, although detection quality and latency were measured in separate evaluations.

own 22

Figure 2: Detection performance of the post-trained 9B model relative to each comparison model on the benchmark. Values are percentage-point differences, not relative percentage improvements.

Against the Opus 4.5 decision agent, the post-trained model improved precision by 7.5%, recall by 12.0%, and F1 by 9.6%. Dollar-weighted recall increased by 35.4%, indicating substantially better coverage of fraudulent transaction value.

Serving with Ray Serve LLM on Anyscale

We served the merged LoRA checkpoint with Ray Serve LLM on Anyscale, using vLLM and an OpenAI-compatible API. To minimize online latency, the serving contract required only a compact risk classification. We deployed the model on L40S GPUs, autoscaling between 2 and 8 replicas based on request load. Ray Serve LLM handled load-aware request routing across replicas while continuously monitoring replica health. As traffic changed, Ray Serve adjusted the number of model replicas, while Anyscale provisioned the underlying GPU capacity needed to run them.

For online production traffic, P50 end-to-end LLM-request latency was 55% lower than Opus 4.5: 0.683 seconds versus 1.515 seconds. P99 was 1.036 seconds versus 3.312 seconds.

own 33

Figure 3: P50 and P99 latency for online serving. 

Owning the Model Also Means Owning Its Lifecycle

Open weights let us choose what to train, which checkpoint to deploy, and when to upgrade. Self-hosting also keeps transaction context out of external model APIs. But fraud patterns change, so maintaining quality requires continued evaluation and retraining—not just a successful checkpoint.

Three disciplines make that control useful:

  • Evaluate domain outcomes: Test candidate checkpoints for both detection quality and fraud-value coverage.

  • Version the whole agent: Track model weights, prompts, and decision policy together.

  • Operate for reliability: Monitor quality and latency under load, with a tested rollback path.

A shared orchestration layer across post-training and serving makes this lifecycle easier to operate as a continuous loop. Running SkyRL and Ray Serve LLM on Anyscale allowed the team to carry the model from post-training into autoscaled production serving while keeping engineering focus on evaluation, model quality, and application performance. 

From Building an Agent to Owning Intelligence

This series began by showing that an LLM agent could strengthen our existing risk stack. Model comparisons established the need for our own evaluation; post-training gave us a way to improve the model itself.

For Coinbase Onramp, owning intelligence means more than owning weights: a benchmark that defines success, a training process that improves decisions, and a serving system we can operate on our terms.





Storie recenti

Informativa sulla responsabilità: il trading di derivati tramite la piattaforma Coinbase Advanced è offerto ai clienti SEE idonei da Coinbase Financial Services Europe Ltd. (licenza CySEC 374/19). Per poter accedere ai derivati, i clienti dovranno superare i nostri controlli di valutazione standard per determinare la loro idoneità e adeguatezza per questo prodotto.