An experiment in cross-market data fusion. A reinforcement learning agent trained to trade Polymarket's 15-minute crypto prediction markets by fusing Binance futures order flow with Polymarket orderbook data.
The thesis: read the "fast" market (Binance) and trade the "slow" market (Polymarket) before the price adjusts.
Note: This represents ~10 hours of paper trading data from a single run on New Year's Eve 2025. The model traded with a $500 fixed position size and $2,000 max exposure (up to 4 concurrent positions).
Results
Metric
Value
Total PnL
$50,195
Return on Exposure
2,510%
Sharpe Ratio
4.13
Profit Factor
1.21
Total Trades
29,176
Win Rate
23.9%
Runtime
~10 hours
Learning Progression
Comparing first 25% vs last 25% of trades:
Phase
Avg PnL/Trade
Win Rate
First 25%
+$1.27
22.5%
Last 25%
+$3.56
25.3%
2.8x improvement in avg PnL per trade. Last 25% of trades generated 52% of total profit.
Limitations: Single 10-hour run. No out-of-sample validation. Results could reflect market regime, not learned behavior. We're sharing the raw data—draw your own conclusions.
Performance by Asset
Asset
PnL
Trades
Win Rate
BTC
+$38,794
8,257
32.6%
ETH
+$9,978
7,859
27.0%
SOL
+$1,752
6,310
16.3%
XRP
-$328
6,750
16.7%
Architecture
LACUNA (v5) uses a temporal PPO architecture:
Temporal Encoder: Sees last 5 states instead of just the present
Asymmetric Actor-Critic: Separate networks for policy and value
Feature Normalization: Stabilizes training across different market conditions
Model Constraints
Fixed position size: $500 per trade
Max exposure: $2,000 (up to 4 concurrent positions)
Fuses data from two sources into an 18-dimensional state:
Category
Features
Momentum
1m/5m/10m returns
Order flow
L1/L5 imbalance, trade flow, CVD acceleration
Microstructure
Spread %, trade intensity, large trade flag
Volatility
5m vol, vol expansion ratio
Position
Has position, side, PnL, time remaining
Regime
Vol regime, trend regime
Training Evolution
Five phases over three days. Each taught us something. Only the last earned a name.
Phase 1: Shaped Rewards (Failed)
Duration: ~52 min | Trades: 1,545 | Result: Policy collapse
Started with micro-bonuses to guide learning:
+0.002 for trading with momentum
+0.001 for larger positions
-0.001 for fighting momentum
What happened: Entropy collapsed from 1.09 → 0.36. The agent learned to game the reward function—collect bonuses while ignoring actual profitability. Buffer showed 90% win rate while real trade win rate was 20%.
Lesson: Reward shaping backfired here. When shaping rewards were gameable and similar magnitude to the real signal, the agent optimized the wrong thing.
Phase 2: Pure Realized PnL
Duration: ~1 hour | Trades: 2,000+ | Result: 55% ROI
Stripped everything back:
Reward ONLY on position close
Increased entropy coefficient (0.05 → 0.10)
Simplified actions (7 → 3)
Smaller buffer (2048 → 512)
Update
Entropy
PnL
Win Rate
1
0.68
$5.20
33.3%
36
1.05
$10.93
21.2%
Win rate settled at 21%—below random (33%)—but profitable. Binary markets have asymmetric payoffs. (Still using probability-based PnL at this point.)