๐Ÿ“Š Fraudalysis Benchmarks

Independent, reproducible benchmark results proving Fraudalysis detection accuracy against industry-standard fraud datasets. Run on Kaggle notebook infrastructure.

Results last verified: 3 October 2026 (AEST) โ€” every figure on this page was checked against the captured results.json from each run. The three Kaggle datasets were last executed on18 September 2026 (synthetic), and3 October 2026 (credit card and Ethereum); run times are taken from the Kaggle API and recorded inprovenance.json. Each score is reproducible from the notebooks in that repository.

Dataset size comparison chart

Results at a Glance

All four datasets below were analysed with identical methodology โ€” a Random Forest classifier (100 trees) run on CPU โ€” so the scores are directly comparable rather than each tuned for its own dataset.

DatasetUnitSizeFraud RateROC-AUC
Metaverse FinancialTransactions78,6008.26%100% detection
Synthetic FinancialTransactions6,362,6200.13%0.9994
Credit Card FraudTransactions284,8070.17%0.9775
Ethereum FraudAddresses9,84122.14%0.9879
Bar chart comparing ROC-AUC across all three independent datasets
Note on units: the first two datasets are analysed per transaction. The Ethereum dataset is analysed per address โ€” each row is one Ethereum address summarised by its aggregate on-chain behaviour. The scores are comparable; the volumes are not the same kind of unit.

Dataset 1: Metaverse Financial Transactions

78,600Transactions Analysed
100%Fraud Detection Rate
< 2sProcessing Time
8.26%Dataset Fraud Rate

Source: Kaggle โ€” faizaniftikharjanjua/metaverse-financial-transactions-dataset

Description: 78,600 blockchain financial transactions from the Open Metaverse, including scam and phishing transactions labelled as high-risk. Contains real blockchain sending/receiving addresses, transaction amounts, behavioural patterns, and verified risk scores.

Transaction Type Breakdown

TypeCount% of TotalFraud Rate
Sale25,04031.86%0%
Purchase24,94031.73%0%
Transfer22,12528.15%0%
Scam3,9495.02%100%
Phishing2,5463.24%100%

Key Findings

  1. 100% detection of scam and phishing transactions โ€” all 6,495 high-risk cases correctly identified by Fraudalysis
  2. New users commit 100% of fraud โ€” veteran and established users showed zero fraud activity across 78,600 transactions
  3. Random purchase pattern = 24.8% fraud rate โ€” users with scattered buying behaviour are significantly higher risk than focused or high-value patterns
  4. Amount-based detection alone is insufficient โ€” average fraudulent (495.35) and legitimate (502.57) transaction amounts are nearly identical
  5. Risk scoring system validated โ€” fraud transactions averaged 97.7 risk score vs 36.3 for legitimate transactions

Fraud by Behavioural Pattern

PatternTotalFraud CasesFraud Rate
Random26,1456,49524.84%
Focused26,03300.00%
High Value26,42200.00%

Fraud by User Age Group

Age GroupTotalFraud CasesFraud Rate
New26,1456,49524.84%
Established26,03300.00%
Veteran26,42200.00%

Fraud by Hour of Day

Top 5 most risky hours for fraudulent activity:

HourFraud RateFraud / Total
00:00 (midnight)9.24%307 / 3,323
22:00 (10 PM)9.13%303 / 3,318
19:00 (7 PM)8.63%281 / 3,255
21:00 (9 PM)8.62%286 / 3,318
17:00 (5 PM)8.53%288 / 3,377
Metaverse fraud by transaction type chart

Dataset 2: Synthetic Financial Datasets For Fraud Detection

6.3MTransactions Analysed
0.9994ROC-AUC Score
21sProcessing Time
0.13%Dataset Fraud Rate

Source: Kaggle โ€” ealaxi/paysim1

Description: 6,362,620 simulated financial transactions covering five transaction types (CASH_OUT, PAYMENT, CASH_IN, TRANSFER, DEBIT) over 742 simulated hours. Built from a real-world mobile money operator's transaction patterns. Widely used as an industry benchmark for fraud detection systems.

Transaction Type Breakdown

TypeCount% of TotalFraud Cases
CASH_OUT2,237,50035.17%4,116
PAYMENT2,151,49533.81%0
CASH_IN1,399,28421.99%0
TRANSFER532,9098.38%4,097
DEBIT41,4320.65%0

Key Findings

  1. Fraud is exclusively in CASH_OUT and TRANSFER transactions โ€” zero fraud found in PAYMENT, CASH_IN, or DEBIT types across all 6.3M transactions
  2. Fraudulent amounts are 8x larger than legitimate โ€” mean fraud amount ($1,467,967) vs legitimate ($178,197), a clear signal for amount-based heuristics
  3. Flagged fraud system validated โ€” only 16 transfers over 200K were flagged, all of which were actual fraud cases
  4. Sender balance before transaction is the #1 predictor โ€” oldbalanceOrg alone accounts for 34.6% of model importance, far ahead of any other feature
  5. Limited time window for detection โ€” data spans only 742 hours (~31 days), meaning fraud detection must work fast on fresh accounts

Machine Learning Model Results

ClassPrecisionRecallF1-ScoreSupport
Legitimate1.000.990.993,285
Fraud0.990.990.991,643

Model: Random Forest (100 trees) โ€” trained on 24,639 samples, tested on 4,928.
Overall Accuracy: 99% | ROC-AUC Score: 0.9994

Top Fraud Predictors (Feature Importance)

RankFeatureImportance
1oldbalanceOrg34.6%
2amount20.3%
3newbalanceOrig17.4%
4type_encoded16.2%
5oldbalanceDest6.0%
6newbalanceDest5.5%

The top 4 features account for 88.5% of the model's predictive power โ€” sender's balance before/after and transaction amount are the dominant signals.

Feature importance chart showing top fraud predictors

Fraud by Transaction Type

TypeTotalFraud CasesFraud Rate
TRANSFER532,9094,0970.77%
CASH_OUT2,237,5004,1160.18%
Synthetic dataset fraud cases by transaction type

Dataset 3: Credit Card Fraud Detection

284,807Transactions Analysed
0.9775ROC-AUC Score
98.8%Fraud Precision
86.7%Fraud Recall

Source: Kaggle โ€” mlg-ulb/creditcardfraud

Description: 284,807 anonymised credit card transactions from a two-day window in September 2013, reduced by the publisher to 28 principal components (V1โ€“V28). Contains 492 confirmed fraud cases โ€” a 0.1727% rate, one of the most heavily imbalanced fraud datasets in common use.

Why this dataset matters: it is the industry-standard credit card fraud benchmark, and at a 1:578 fraud-to-legitimate ratio it punishes naive modelling. Any accuracy figure above 99.9% here is meaningless โ€” a model that flags nothing scores 99.83%. We therefore report ROC-AUC alongside precision and recall.

Key Findings

  1. ROC-AUC of 0.9775 with 98.8% precision โ€” when this model flags a transaction as fraudulent it is almost always correct
  2. Fraudulent transactions are mostly small โ€” median 9.25 vs 22.00 for legitimate, despite fraud having the higher mean (122.21 vs 88.29)
  3. Legacy fraud was more aggressive in value terms โ€” the largest fraudulent transaction was 2,125.87 against a legitimate maximum of 25,691.16
  4. V10 and V14 dominate the signal โ€” together accounting for roughly 30% of the model's predictive power

Amount Pattern: Why Averages Mislead

The mean and the median tell opposite stories here, and that matters for anyone building threshold-based rules. Looking only at averages would suggest fraudulent transactions are larger than legitimate ones. In truth, fraudulent card fraud clusters into many small transactions, and the higher mean is pulled up by a small number of large cases.

Credit card fraud amount distribution showing fraud transactions cluster at low values despite a higher mean

Top Fraud Predictors

Feature importance chart for the credit card fraud benchmark

Detection Outcomes

Confusion matrices comparing credit card and Ethereum detection results

On the credit card test set, the model correctly identified 85 of 98 fraudulent transactions (13 missed) while raising only 1 false alarm across 296 legitimate transactions.


Dataset 4: Ethereum Fraud Detection

9,841Addresses Analysed
0.9879ROC-AUC Score
97.1%Fraud Precision
85.3%Fraud Recall

Source: Kaggle โ€” vagifa/ethereum-frauddetection-dataset

Description: 9,841 Ethereum addresses flagged for fraudulent activity, characterised by 47 behavioural features โ€” transaction counts, timing between transactions, unique counterparties, total Ether moved, ERC20 token activity, and contract-creation counts. 2,179 addresses (22.14%) are labelled fraudulent.

What this dataset actually measures. Each row is one address summarised by its aggregate on-chain behaviour โ€” not one individual transaction. The headline number is therefore addresses detected, not transactions detected. This is real recorded blockchain data with real transaction history behind each row, but the unit is an account and we label it as such rather than inflating it into a transaction count.

Key Findings

  1. ROC-AUC of 0.9879 on real blockchain data โ€” the same model and methodology that scored 0.9994 on synthetic transactions
  2. Fraudulent addresses move far less Ether โ€” legitimate addresses sent a mean of 13,025.75, fraudulent addresses only 87.37
  3. Timing is the strongest single signal โ€” "time between first and last transaction" ranked as the top predictor
  4. Fraud is associated with low, dispersed activity โ€” fraudulent addresses averaged 5.17 transactions sent against 147.43 for legitimate ones, across far fewer unique counterparties

Top Fraud Predictors

The predictors split cleanly into two families: on-chain value and flow (how much Ether moved, how it was received) and behavioural timing (how long an address has been active, how regularly it transacts).

Feature importance chart for the Ethereum fraud benchmark

Detection Outcomes

Across the 1,969 held-out test addresses the model caught 372 of 436 fraudulent addresses (85.3%), missing 64 and raising 11 false alarms against 1,522 legitimate addresses.


How We Verified These Numbers

Benchmark pages are easy to publish and hard to trust, so we publish the method alongside the result โ€” including the parts that made our own numbers worse.

We Do Not Claim GPU Acceleration We Did Not Use

An earlier version of our synthetic benchmark script advertised itself as "GPU-accelerated" in both its description and its console output. It was not. The estimator is scikit-learn Random Forest, which runs on CPU only, and the script's gpu_accelerated field was a hardcoded false with a comment explaining it โ€” meaning nobody had ever actually checked.

We fixed this by probing the hardware directly and recording what is really there. Every benchmark now reports three separate fields instead of one claim:

These benchmarks ran on Kaggle instances that genuinely did have an NVIDIA Tesla T4 available, and the published results.json files record gpu_present: true alongside gpu_accelerated: false. Both statements are true at once: the hardware was there, and we did not use it. You can verify this yourself in the raw JSON.

Why not just use the GPU? scikit-learn's Random Forest has no GPU implementation. Switching to XGBoost with device="cuda" would mean benchmarking a different algorithm and breaking comparability across all three datasets. On a 6.3M-row dataset the CPU run took 21 seconds โ€” GPU acceleration would have bought speed we did not need, at the cost of a methodology we could not honestly compare.

We Removed a Score That Made Us Look Better

Our first working version of the Ethereum benchmark reported a ROC-AUC of 0.9971. We did not publish it.

829 of the 9,841 addresses (8.4%) have a blank cell in every ERC20 column โ€” they had never transacted with an ERC20 token at all. Our initial code filled those gaps with the column median, which handed the model plausible-looking token volumes for addresses that had none. The score went up. The data was fiction.

"No ERC20 activity" is a real, meaningful zero โ€” not a missing value to be guessed at. Filling those cells with 0 instead dropped the score to 0.9885, and the run on Kaggle's actual runtime returned 0.9879.

Chart comparing the fabricated 0.9971 score against the honest 0.9879 result
The principle: a benchmark that has been quietly tuned to look better is worth nothing to the person relying on it. We would rather publish 0.9879 and show you exactly how we got there.

Reproduce It Yourself

Every number on this page comes from a script you can read and re-run. Nothing is transcribed by hand, and the charts are generated directly from the captured result files.


Methodology

All tests were conducted using the Fraudalysis internal analysis engine running on Kaggle notebook infrastructure (NVIDIA T4 GPU available; CPU-based scikit-learn estimators used, with no GPU acceleration). The analysis script loaded the raw CSV dataset, classified each record based on behavioural patterns, transaction type, and risk score, then cross-referenced against the dataset's labelled anomaly field.

For all four datasets the classifier is a Random Forest ensemble of 100 trees with random_state=42, evaluated on a stratified 20% hold-out split. Keeping the estimator and seed identical across datasets is deliberate: it means the scores reflect how each dataset differs, not how each script was tuned.

The full analysis code and raw results are available on GitHub for independent verification and reproduction.

What's Next

We continue to test against additional datasets. Candidate sources under evaluation include:

Results will be published here as each benchmark is completed, with the same methodology and the same disclosure of any figure we chose not to publish. Stay tuned.