AI-written summary of reporting by CryptoSlate. No human editor reviewed this. AI can misread or omit facts — read the original, linked below.
ABSTRACT

Coinbase reported on Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, under an unchanged decision policy. The evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions.

Coinbase says newer AI models caught less fraud in Onramp screening replay

Coinbase says newer AI models caught less fraud in Onramp screening replay

Coinbase reported on Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, under an unchanged decision policy. The evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions.

Context

The cohort covered nine weeks before the company's risk agent rolled out, retaining all matured fraud cases while sampling legitimate traffic. Coinbase said each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions, which it said isolated the decision model's behavior within that setup rather than comparing redesigned screening systems.

Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). It said every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Sonnet's recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points, while Opus's recall declined 0.8 points, according to the company. Both newer models also had lower precision, meaning a smaller share of transactions they classified as fraud were actually fraudulent.

GPT's precision rose 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points, the company said. Coinbase said its fraud flags were more accurate while more fraud cases and value escaped detection in the replay, and that the replay does not establish customer losses from deploying those versions. Coinbase also said it could identify the regressions without establishing their cause.

Coinbase's earlier online experiment compared adding selective LLM review with the existing models and rules alone. It said that agent-enabled flow recorded 30% fewer fraudulent transactions and 22% less fraud value, and that it did not compare newer model versions. The related SR-Fraud payment-fraud study first appeared Sept. 23 and was revised Sept. 30, before the October blogs, according to the source.

In its Oct. 8 disclosure, Coinbase reported that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics, with F1 up 9.6 percentage points and dollar-weighted recall up 35.4 points. It said it specialized the model using historical fraud outcomes and deterministic rewards balancing fraudulent and legitimate examples. Separately, production measurements put median end-to-end LLM-request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5, a 55% relative reduction; the source states the faster inference and stronger benchmark detection came from different evaluations.

Gaps & Unknowns
  • The source does not state the absolute precision, recall, F1 or dollar-weighted recall values for any of the models; only the changes between versions are given.
  • The source does not name the developers of the Opus, Sonnet or GPT model families and does not state whether they were asked to respond.
  • The source does not state the cause of the reported declines, beyond Coinbase saying it could identify the regressions without establishing their cause.
  • The source does not state the calendar dates covered by the nine-week cohort or the dates of the model upgrades.
  • The source does not identify the four fraud-detection metrics used in the Qwen3.5-9B comparison, beyond F1 and dollar-weighted recall.
  • The source does not state the years of the Sept. 23 and Sept. 30 dates given for the related SR-Fraud study.
Sources & Further Reading
  1. CryptoSlate — original

Read the original at CryptoSlate

Related Coverage