Is a Challenge "Expected-Value Positive Just by Buying It"? I Built a Coin-Toss EA and Backtested It 16 Times [Analysis]
※ This article is an analysis (opinion piece), not investment advice. The backtests are validation against historical data and do not guarantee future performance. This is not a recommendation to trade with random entries, and the text touches on the possibility of running into each firm's terms of service. Firm-specific rules are updated frequently, so always check each firm's official site. (Last updated: August 31, 2026)
The question: is a challenge a "good deal" against its fee?
Last time, we showed with an exact solution that even a coin toss clears Phase 1 around half the time. That naturally leads to the next question.
If you pay a ¥108,800 fee and the expected value beats that, shouldn't you just mechanically keep buying?
There's also one trick involved: placing SL/TP at 40x the spread. Doing this fixes the per-trade cost at "one spread's worth ÷ 40 = 2.5% of the stake" every time, and the absolute value of the currency pair drops out of the equation.
Is this idea actually correct? I wrote a dedicated EA, then built the same logic into the production EA (ELDRA) as well, and ran a total of 24 backtests.
Here are the 3 conclusions up front:
- The theory checked out completely. The theoretical win rate was 48.75% against a measured 48.71%. Just as the formula predicted, changing the risk-reward ratio doesn't move the per-trade expected value at all
- The expected value was positive. Solved from the measured trade distribution, a single challenge's expected value came out to +¥119,000 against a ¥108,800 fee (95% confidence interval +¥18,000 to +¥381,000)
- But the bottleneck isn't the math. Gaps, terms of service, and "can you actually withdraw once funded" remain
One thing up front: the numbers through Section 9 are all under a "once a day" restriction. That's an experimental condition meant to isolate the per-trade economics — it's not a recommended setting. In Section 10, once that restriction is removed, the time to resolve a single challenge shrinks from 139 days to 7 days, with no drop in the pass rate. If you're in a hurry, start reading from Section 10.
1. What does "spread x M" actually fix in place?
Placing SL and TP at an equal distance (M x spread) from the entry price gives you this:
| Move size | Probability | |
|---|---|---|
| Win | +M x spread | p = (M-1) / 2M |
| Loss | -M x spread | 1-p |
The payoff is symmetric, yet the win rate falls below 50% — that's the crux of it. A buy enters at the Ask and the TP fires when the Bid reaches it, so winning only needs (M+1) spreads of movement while losing only needs (M-1) spreads. This asymmetry is what "the spread you're paying" actually is.
At M=40, p = 39/80 = 48.75%. The expected value per trade is
M*s x (p - (1-p)) = M*s x (-1/M) = -s
Exactly one spread's worth. No information about the currency pair survives in there at all. The idea checks out.
Changing the risk-reward ratio doesn't move the expected value one bit
The natural next thought is "so what if I set the risk-reward to 1:2, 1:3, or 1:4?" — but that turns out to be mathematically meaningless. Setting TP to RR times the distance makes the win rate p = (M-1) / (M(RR+1)), and when you compute the expected value, the RR cancels out cleanly and disappears.
M*s x (p*RR - (1-p)) = -s
Whether RR is 1 or 4, the expected value per trade is always exactly -1 spread. What RR affects isn't the expected value — it's the path via the number of trades needed. Raising RR speeds up resolution, and as a result cuts down the total spread paid (more on this later).
2. The industry frames it the same way
This view isn't new — outside Japan, it's commonly discussed as a "fee cost per dollar of drawdown" metric. For a $100K account with a $500 fee and a $10,000 max DD, that's $0.05 per DD dollar (5.0%).
Why divide by the drawdown? Because for a fair bet, the total profit you can expect from a funded account is exactly equal to the max DD.
This follows from the martingale (optional stopping theorem for fair bets). Suppose you repeat a cycle of "withdraw and reset to the starting balance once the balance hits +W%" until you fail. The probability of one cycle succeeding is p = DD/(W+DD), the expected number of successful cycles is p/(1-p) = DD/W, and since each one withdraws W%,
Expected total profit = W x DD/W = DD
No matter how you choose the withdrawal step W, the expected total extraction is exactly the max DD. The amount the firm is putting at risk is exactly what comes back out — an intuitive result.
For Fintokei ProTrader, the max DD is 10% and the split is 80%, so the expected value of a single funded account is 8% of the account. For the ¥20M Sapphire tier, that's ¥1.6M. The fee is ¥108,800 (0.544% of the account).
3. The model's expected value
All that's left is to chain the gates.
EV = P(Phase1) x P(Phase2) x 0.8 x max DD - fee
Here's the exact solution, solved as a linear system from the absorbing Markov chain (Fintokei Sapphire, ¥20M, 8% → 6% target, 10% max DD, 80% split, fee 0.544%).
| SL width | Risk per trade | Win rate | Phase1 | Phase2 | Funded expected value | EV |
|---|---|---|---|---|---|---|
| x20 | 1% | 47.50% | 34.0% | 43.5% | 3.90% | +¥33,000 |
| x40 | 1% | 48.75% | 44.4% | 52.9% | 7.67% | +¥180,000 |
| x40 | 4% | 48.75% | 57.0% | 57.0% | 10.60% | +¥442,000 |
| x80 | 4% | 49.38% | 58.5% | 58.5% | 11.27% | +¥508,000 |
On the model, it's clearly positive. At x40 / 4% risk, that's roughly 4x the ¥108,800 fee.
And there's a counterintuitive result: the bigger the risk per trade, the better. The reason is simple — total cost is inversely proportional to (risk x SL width). The number of trades required shrinks in inverse proportion to the square of the risk, so betting bigger means paying less total spread. Raising RR lowers total cost for the same reason (at x40 / 1% risk, RR 1:1 needs 127 trades vs. RR 1:4 needs only 38).
That's the theory so far. The question is whether this holds up in reality.
4. I wrote a dedicated verification EA
Since ELDRA is a production EA that's being distributed, I left it untouched and wrote a brand-new minimal EA purely for verification. With no other features to muddy the results, it's a clean test of the hypothesis.
- Entry direction is a coin toss via
MathRand()(seed varies the run) - SL = the recent median spread x M, TP = SL x RR
- Lot size is backed out via
OrderCalcProfit()so that the loss at SL equals R% of the reference balance - No new orders while the spread exceeds 1.5x the median
- Capped at once a day (※ as covered below, this restriction turned out to be unnecessary)
The reason for capping it at once a day was to isolate the per-trade economics alone. Trading many times a day drags the daily drawdown wall into the picture, making it impossible to separate the spread effect from the daily-DD effect. The plan was to measure the bare single-trade case first, then raise the frequency afterward.
As it turned out, this restriction was just pointlessly killing the turnover rate, and it gets removed in Section 10. Read the day counts through Section 9 with that restriction in mind.
The reason for using the median spread was to avoid stepping on the Monday-open spike. In the first run, a trade got opened during a momentary 160-point spread and melted 60% of the account in one shot.
Conditions matched Fintokei ProTrader's Sapphire tier (¥20M, 100% real ticks, September 2025 - August 2026, XAUUSD). Since Fintokei ProTrader's max DD and daily DD are both static and balance-based, floating losses can be ignored and judgment can be made purely on realized P&L.
5. Measured results: the theory landed exactly on target
3 seeds each across RR settings, 12 runs total, plus 4 currency-pair comparison runs.
| RR | Trades | Measured win rate | Theoretical win rate | Avg. win | Avg. loss |
|---|---|---|---|---|---|
| 1:1 | 735 | 48.71% | 48.75% | +0.9905% | -1.0078% |
| 1:2 | 648 | 34.72% | 32.50% | +1.9914% | -1.0114% |
| 1:3 | 589 | 23.26% | 24.38% | +2.9436% | -1.0174% |
| 1:4 | 546 | 22.34% | 19.50% | +3.8701% | -1.0168% |
At RR 1:1, measured came to 48.71% against a theoretical 48.75%. A 0.04-point difference. The formula p=(M-1)/2M held exactly against real data.
I also got a breakdown of costs (per trade, relative to the account):
| Item | Value |
|---|---|
| Mark-to-market P&L | -0.0689% |
| Commission | -0.0016% |
| Swap | -0.0016% |
| Total | -0.0721% |
Commission and swap combined came to just 0.0032%, essentially negligible — that was surprising, given an average holding time of 12.6 hours, which crosses overnight.
On the other hand, average win of 0.9905% against average loss of 1.0078% means the loss side is slightly bigger. That's the fill difference between limit and stop orders (stop-side slippage) — a real cost the theory doesn't account for.
6. But this test has a decisive limitation
This is the most important part of the article.
Calculating the measurement precision of the per-trade expected value gives this:
| RR | Trades | Per-trade EV | 95% CI | Verdict |
|---|---|---|---|---|
| 1:1 | 735 | -0.0345% | -0.108 to +0.039% | Not inconsistent with theory |
| 1:2 | 648 | +0.0313% | -0.080 to +0.143% | Not inconsistent with theory |
| 1:3 | 589 | -0.0961% | -0.232 to +0.040% | Not inconsistent with theory |
| 1:4 | 546 | +0.0751% | -0.098 to +0.248% | Not inconsistent with theory |
Every confidence interval straddles zero. RR 1:2 and 1:4 have point estimates that look "positive," but that's noise. As long as you're paying a cost on a fair coin, a positive expected value is mathematically impossible.
The reason is clear: the effect being measured (-0.025%) is only 1/40th the size of a single trade's own noise (±1%).
Detecting -0.0250% at 95% confidence requires about 6,532 trades. A year of backtesting only produced 735. That's short by a factor of 9.
This isn't "the backtest was done badly" — it's that the per-trade expected value isn't a quantity you can measure with a backtest. At once-a-day pace, you'd need 26 years of data.
However, while "the per-trade expected value" can't be measured, "a single challenge's expected value" can be. That's because a tiny difference gets amplified into the binary outcome of "does it hit the wall, or does it reach the target." That estimate is done in Section 11.
In cases like this, the exact solution becomes the more reliable source of information. The result that the expected value is fixed at -1 spread is a structural consequence that holds as long as prices follow a martingale — it barely requires any assumption about the price distribution.
Meanwhile, the challenge pass rate itself agreed well with theory, when measured using only non-overlapping (independent) windows.
| Risk per trade | Measured (independent windows) | 95% CI | Theoretical |
|---|---|---|---|
| 1% | 33.3% (n=6) | 0 to 71% | 44.4% |
| 2% | 45.5% (n=33) | 28.5 to 62.4% | 50.0% |
| 4% | 47.1% (n=87) | 36.6 to 57.6% | 57.0% |
※ Measuring with overlapping windows collapses the effective sample size down to just a few, producing completely untrustworthy numbers. The author got caught by this along the way and initially — and incorrectly — concluded "there's serial correlation." Running a variance-ratio test showed no significant correlation; it was simply an artifact of overlapping windows.
7. "Doesn't matter which pair" is only half true
This became clear once tested. The cost ratio really is independent of the pair, but the trading frequency differs by an order of magnitude.
| Symbol | Trades/year | Per trade |
|---|---|---|
| XAUUSD | 247 | 1.5 days |
| EURUSD | 63 | 5.7 days |
| USDJPY | 24 | 13.5 days |
| GBPJPY | 10 | 32.3 days |
Under the same "spread x 40" setup, the time to resolution differs by a factor of 20. That's because the ratio of spread to volatility differs from pair to pair. GBPJPY averages 32 days per trade, meaning it would take decades to complete a single challenge — effectively unusable.
The expected-value formula doesn't depend on the pair, but the time it takes to realize that formula depends on the pair heavily. If you're optimizing for time efficiency, you're limited to symbols with high volatility relative to their spread (gold, in this test).
8. Where's the weakest assumption?
The one part of the model that hasn't been backed by real data is the assumption that "once funded, you can extract the full max-DD amount as expected." If that breaks down, the sign of the expected value flips.
Checking overseas statistics, the numbers were split right down the middle.
| Source | Funded → reaches payout |
|---|---|
| Track360 | 45% |
| QuantVPS (TopStep) | 7% |
A 6x gap. But this isn't a contradiction — it's explained by the difference in drawdown method. TopStep is a futures firm with a trailing DD, where the lower wall rises along with your profit, so a cycle of "grow to +W%, withdraw, start again from zero" doesn't work. With a static-DD firm, on the other hand, the reset wall position doesn't move even after a withdrawal.
Fintokei ProTrader has a static DD, so in theory it sits in Track360's world (45%). The "probability of reaching the first payout from a funded account" that my model produces is also 40-60%, which is consistent with that number.
Run this strategy at a trailing-DD firm, and the expected extraction falls well short of the max DD, flipping the expected value negative. The extent to which the difference in DD method matters this much was a discovery even for the author.
9. What "there's an edge if you commit to it" actually means: how many runs to realize it
Even if the expected value is positive, that's an average. I calculated how many runs it takes before it actually lands in your hand.
At x40 / 2% risk, a single challenge is a bet that looks like this:
| Probability | Outcome | |
|---|---|---|
| Reaches funded | 28.9% | → of which 81.2% get even one payout |
| Reaches a payout | 23.4% | Hit |
| Fails | 76.6% | The ¥108,800 fee disappears |
3 out of 4 runs lose the fee entirely. The expected value per run is +¥290,000, but the standard deviation is ¥1.034M — 4x the expected value.
From there, you can derive how many runs you'd need before cumulative results are positive with 95% confidence.
| Setting | Funded rate | Reaches payout | Per-run EV | Std. dev. | Runs needed | Total fees needed |
|---|---|---|---|---|---|---|
| x40 / R1% | 23.5% | 20.8% | +¥180,000 | ¥820,000 | 56 | ¥6.1M |
| x40 / R2% | 28.9% | 23.4% | +¥290,000 | ¥1.034M | 34 | ¥3.75M |
| x40 / R4% | 32.5% | 18.5% | +¥442,000 | ¥1.506M | 31 | ¥3.42M |
| x80 / R2% | 31.7% | 26.1% | +¥363,000 | ¥1.153M | 27 | ¥2.98M |
At minimum, 27-34 runs, for a total fee outlay of ¥3M-¥3.75M. That's the concrete substance behind "there's an edge if you commit to it."
You also need to be able to withstand a losing streak. The probability of failing to reach a payout on 10 straight runs is 6.9% (happens more than once in 20 tries), and the fees paid during that stretch would be ¥1.09M. A positive expected value and being able to withstand that path financially and psychologically are separate problems.
And trying to rack up enough runs runs straight into allocation caps. FTMO caps a single trader's total at $400,000, so you can't hold dozens of accounts simultaneously.
That said, "how many days per run" changes drastically depending on the setting. Up to this point we've been capped at once a day, but raising the number of daily attempts up to the daily loss limit sped up resolution by an order of magnitude. Measured directly in the next section.
10. Raising the number of daily attempts up to "just short of using the whole daily allowance"
Up to now, I fixed the trade count at once a day to isolate the per-trade economics. Here, that restriction gets removed.
If the daily loss cap is 5% and the risk is 2%, you should be able to keep trading as long as taking one more loss would still stay inside the cap. And under this rule, the daily-DD wall never triggers, by construction. There's no benefit to leaving the allowance unused — it's just throwing away time.
I built this as a module into the production EA, ELDRA, and measured it on XAUUSD (an off-by-default add-on feature, MQL5 only).
Result: resolution up to 35x faster
| Setting | Trades/day | Time to resolve one challenge | Phase 1 pass rate | Independent samples |
|---|---|---|---|---|
| Once/day, 1% risk (previous sections) | 1.0 | 139 days | 33.3% | 3 |
| Up to daily allowance, 1% risk | 6.5 | 19 days | 38.9% | 36 |
| Up to daily allowance, 2% risk | 4.2 | 9 days | 30.0% | 50 |
| Up to daily allowance, 4% risk | 3.0 | 4 days | 41.6% | 89 |
139 days became 4-19 days, and the pass rate didn't drop. The once-a-day restriction had no purpose at all — it was just throwing away time.
This is also consistent from an expected-value standpoint. Since the per-trade expected value is fixed at one spread's worth, trading more times in the same day doesn't change the per-trade cost. The distance to the wall doesn't change either. All that changes is "the real time it takes to reach the wall."
Was the daily drawdown actually respected?
| Setting | Worst daily P&L | Days exceeding 5% daily |
|---|---|---|
| 1% risk | -4.51% | 0 days |
| 2% risk | -4.51% | 0 days |
| 4% risk | -9.64% | 4 days |
At 1% and 2% risk, the 5% daily limit was never touched even once, across 497 and 306 days respectively. The "stop just short of using the whole allowance" rule worked exactly as designed.
The 4% risk case breaking down is because a single gap loss can vastly exceed the intended 4%. More on that next.
The one remaining hole: gaps jump the stop-loss
3.3-3.5% of all losses exceeded the intended risk by more than 10%. The worst case was on Monday, April 13, 2026: a single trade at -8.542%, or 4.27x the intended 2%. A weekend gap jumped the stop-loss.
This is the one place where the EV model's assumption (loss is capped at risk R%) breaks down, and it can't be prevented by managing the daily allowance. The allowance can stop "orders about to be placed," but it can't stop a gap on a position that's already open.
I turned on ELDRA's weekend-close feature and retested, but it didn't help. The reason is obvious once you look at the day-of-week distribution.
| Day-of-week distribution of losses exceeding intended risk | Worst | |
|---|---|---|
| Weekend close OFF | Mon 8, Thu 9, Tue 3, Fri 3, Wed 1 | -3.67% (Mon) |
| Weekend close ON | Mon 8, Thu 8, Wed 8, Fri 3, Tue 2 | -5.04% (Mon) |
Thursday shows just as many as Monday. It's not only weekend gaps — slippage from economic indicators and central bank events concentrated on Thursdays happens at the same scale. Closing over the weekend only plugs half the hole.
Final assessment of time efficiency
In the previous section, we calculated "you need 43 runs to reach cumulative positive results with 95% confidence." Since the days-per-run figure has changed, let's redo the required number of years.
| Setting | One cycle (Ph1+Ph2) | Runs per year | Years needed |
|---|---|---|---|
| Once/day, 1% risk | 243 days | 1.0 | 43 years |
| Up to daily allowance, 1% risk | 33 days | 7.5 | 5.7 years |
| Up to daily allowance, 4% risk | 7 days | 35.7 | 1.2 years |
43 years became 1.2 years. The "on the order of several years" estimate written in the previous section was skewed by the unnecessary once-a-day restriction, and I'm correcting it here.
That said, the 4% risk setting punched through the daily DD on 4 separate days (from gaps). Taking time efficiency costs you gap resilience — that's the trade-off here, and in real operation, something around 2% risk looks like the practical sweet spot.
11. The final verdict: is there an expected value if you run it mechanically?
Section 6 said "a year of backtesting can't measure the per-trade expected value." But the challenge pass-rate expected value can be measured. Measuring a tiny per-trade difference of -0.025% is impossible, but once that gets converted into the binary outcome of "does it hit the wall, or does it reach the target," the signal becomes large enough.
So I switched to a higher-power estimation method. Instead of counting the number of windows (36-89), I fed the entire measured trade-level P&L distribution (877-3,205 trades, including gaps, fees, and swap) into an absorbing Markov chain to solve exactly for the probability of reaching the wall, then produced confidence intervals via 120 bootstrap resamples (with replacement) of the trades. The absence of serial correlation was already confirmed with a variance-ratio test (0.71-0.92, not significant), so the IID assumption is reasonable.
Since this model doesn't include the daily DD, I corrected it using the ratio against the window-based measurement (which does include daily DD).
| Setting | Phase1 | Phase2 | Lost to daily DD | EV per run | 95% CI |
|---|---|---|---|---|---|
| 1% risk (6.5x/day) | 38.9% | 47.9% | -1.6pt | +¥68,000 | -¥27,000 to +¥454,000 |
| 2% risk (4.2x/day) | 30.0% | 35.7% | -7.4pt | -¥34,000 | -¥80,000 to +¥75,000 |
| 4% risk (3.0x/day) | 41.6% | 46.3% | -11.8pt | +¥119,000 | +¥18,000 to +¥381,000 |
| 2% risk + weekend close | 34.4% | 40.8% | -6.9pt | +¥7,000 | -¥59,000 to +¥164,000 |
Only the 4% risk case has a confidence interval that sits entirely on the positive side. That's +¥119,000 per run against a ¥108,800 fee — roughly a 2.1x return. At 35.7 runs a year, that works out to +¥4.25M per year.
But this confidence interval only reflects the variance in the trade distribution itself. The following three things aren't included at all.
- Whether you can actually withdraw once funded. The model relies on the result of the optional stopping theorem, that "you can extract the full max-DD amount." Payout frequency limits, minimum trading days, consistency rules, and firm discretion are all unverified
- Terms-of-service risk. There's no guarantee random entries would be treated as "genuine trading"
- One year, one symbol. This only looked at XAUUSD from September 2025 to August 2026. If the market regime changes, the distribution changes too
It's also worth noting that the best-performing 4% risk case loses a full 11.8 points to the daily DD. That's the portion where the stop-loss got jumped by a gap. The higher the risk, the more a single incident becomes fatal, and in the actual measurements it punched through the 5% daily limit on 4 separate days.
The answer
Mathematically, the expected value is positive. But that positive result only holds up through "the relationship between the wall and the fee" — beyond that, "can you actually get paid out" sits outside the math.
Section 6 said "the per-trade expected value can't be measured with a backtest." But the per-challenge expected value can be measured, and the answer was positive. That's a firmer conclusion than anything stated in earlier sections.
Still, the reason I can't write "run it mechanically and you'll profit" is that between the expected value being positive and that expected value actually landing in a bank account, there are 3 layers that math can't reach.
12. Running multiple accounts in parallel: what matters is "independence," not "speed"
Everything up to this point was about a single account. Anyone actually running this kind of operation runs multiple accounts in parallel, so let's work that out too. Assume 12 accounts.
Turnover rate directly becomes your coupon-redemption capacity
Discount coupons have expiration dates. If every slot you have is already occupied, a coupon showing up just gets skipped. The rate at which open slots free up directly becomes the number of coupons you can actually use.
| Setting | Days per run | Free slots per month | Runs per year |
|---|---|---|---|
| Once/day, 1% risk | 243 days | 1.5 | 12 |
| Up to daily allowance, 1% risk | 33 days | 10.9 | 91 |
| Up to daily allowance, 2% risk | 16 days | 22.5 | 188 |
| Up to daily allowance, 4% risk | 7 days | 51.4 | 429 |
With only 1.5 slots freeing up a month, of course you'd miss coupons. Just switching to a setting that uses up the daily allowance takes that to 10-51 a month. The "trades per day" from the previous section wasn't just about speed — it also determines how many coupons you can actually pick up.
Scaling to 12 accounts doesn't reduce variance if the setting is the same
This is the main point here. Even with the same expected value, the spread of outcomes is completely different depending on the correlation between accounts.
Running 188 runs a year (16-day cycle, 12 accounts) over one year:
| Setup | Expected value | Std. dev. | Probability of finishing positive |
|---|---|---|---|
| Different random seed per account | +¥22.37M | ¥11.07M | 98% |
| Same setting on every account | +¥22.37M | ¥151.72M | 56% |
The expected value is exactly the same, but the probability of finishing positive is 98% vs. 56%. Put the same EA with the same settings on 12 accounts, and the correlation is nearly 1 — adding accounts doesn't cut the variance by a factor of 12, it stays piled up at 12x. The fact of "holding 12 accounts" becomes statistically meaningless.
I think this is the most commonly misunderstood point in multi-account operations. The point of adding more accounts isn't capital — it's diversification, and diversification only comes from independence.
And a correlation of 1 is also dangerous under the terms of service
Setting statistics aside, identically configured multiple accounts run straight into copy-trade detection. There are reported cases where IP fingerprinting and millisecond-level execution matching flag simultaneous fills from the same IP and same lot size within 10 milliseconds, causing all accounts' payouts to be denied at once (details on detection and penalties).
The coin-toss method has a structural advantage here. Simply using a different seed per account makes the trading on all 12 accounts completely uncorrelated. No need for tricks to stagger execution — it's independent both statistically and under the terms of service. Trying to do the same thing with discretionary trading or a logic-based EA requires deliberately scattering parameters, which dilutes the edge in the process.
The next wall isn't "time" — it's "capital" and "combined caps"
Raising turnover sends annual fees soaring (¥20.4M a year on a 16-day cycle, ¥46.63M a year on a 7-day cycle). The working capital tied up at any given moment is only 12 runs' worth, i.e. ¥1.31M, but whether this actually works depends on whether payouts actually land — as noted in Section 8, that's the one unverified piece here.
Combined caps also come into view.
| Firm | Combined cap |
|---|---|
| FundingPips | $2M |
| FTMO | $400K |
| Hantec Trader | $400K |
| Fundora | ¥60M |
Concentrating on a single firm runs into that cap, so raising turnover requires spreading across more firms.
13. So, in the end, is it worth doing?
The math side checks out. The measured results, to the extent they can be measured, matched theory. Even so, I can't write "keep buying mechanically and you'll win." That's because the bottleneck isn't the math.
Checking the terms-of-service side, Fintokei's list of banned practices covers 4 things: arbitrage, copy trading, HFT/tick scalping, and gap trading — a once-a-day random entry falls under none of them. Self-built EAs are explicitly allowed. That said, every firm has a catch-all clause letting them deny payouts on the grounds that trading lacks "genuine substance," and there's no guarantee a random entry would be seen as "legitimate trading." That's an area decided by operational discretion, not the wording of the terms.
Allocation caps also come into play. FTMO caps a single trader's total at $400,000. The approach of running multiple accounts to lean on the law of large numbers (Section 12) is workable in principle, but concentrating on one firm hits the cap, so the more accounts you add, the more firms you need to spread across.
And the variance is extreme. At 4% risk, the max DD of 10% only allows for 2.5 losses' worth of room — 3 in a row ends it. Even if the expected value is positive, realizing it only happens after dozens of runs. And just lining up accounts doesn't reduce variance by itself — as shown in Section 12, if accounts aren't independent, 12 accounts are no different from 1 in terms of the spread of outcomes. Whether you can keep paying the ¥108,800 fee, dozens of times over, through the losing streaks along the way, is a separate problem.
Here's my conclusion:
The structure of "a challenge being cheap relative to its fee" is real, and the expected value solved from the measured trade distribution came out positive too (+¥119,000 per run, 95% CI +¥18,000 to +¥381,000). But that's limited to firms with a static DD, and the bottleneck isn't the math — it's gap resilience, the terms of service, and "can you actually withdraw once funded."
Put another way, this isn't "an arbitrage that pays off the moment you find it" — it's "a thin edge that only the people who can keep at it for a long time can capture." The standard deviation per run is 4x the expected value, so 10 or 20 runs proves nothing. This connects to the same structure we saw in the previous article — that "the consistency rule punishes concentration in time" — and it's fair to say that the prop-firm game as a whole is deliberately designed to make short-term results meaningless.
The previous article said "the average trader is losing to a coin toss." What this one reveals is the flip side of that. If a coin toss has any room to win at all, it's not because it's calling the market right — it's because it isn't doing anything extra.
Methodology
- Theoretical values: the challenge formulated as an absorbing Markov chain, expanded breadth-first over only the reachable states, solved as an exact solution via the linear system (I-A)f = c. Confirmed to match the analytical solution
DD / (target + DD)to 1e-16 precision under zero-fee conditions - Measured: a dedicated verification EA newly written in MQL5, run 16 times in the MT5 Strategy Tester (XAUUSDp and others, M1, 100% real ticks, 2025.09.01-2026.08.28, starting margin ¥20M)
- Section 10 built the same logic into the production EA (ELDRA) as a module and ran 8 more times on XAUUSD (24 runs total). The daily allowance check uses "the balance at the previous day's final tick" as its reference (recalculating on the new day's first tick would let the settlement loss on a position carried overnight leak out of that day's allowance)
- The EV in Section 11 fed the measured full trade-level P&L distribution (877-3,205 trades, including gaps/fees/swap) into an absorbing Markov chain to solve exactly for the probability of reaching the wall, then produced confidence intervals via 120 bootstrap resamples of the trades (with replacement). Since the model doesn't include daily DD, it's corrected using the ratio against the window-based measurement
- The multi-account figures in Section 12 are theoretical values synthesized by varying the correlation coefficient, based on the per-run expected value and standard deviation
- P&L was taken from the difference in the balance series, so it fully includes commission and swap
- The challenge pass rate is aggregated using only non-overlapping (independent) windows. Overlapping windows drastically reduce the effective sample size and aren't reliable
FAQ
Q. Is there a specific reason for using "40" in spread x 40?
Since the cost ratio is just 1/M, a larger M is always better. But raising M stretches out the time to resolve a single trade by the square of M (due to diffusion). x40 landed in a practical range — about 1.5 days per trade on gold. Going to x80 or beyond improves the expected value, but a single challenge can then take years.
Q. Which risk-reward ratio is actually best?
From an expected-value standpoint, they're all completely indifferent (mathematically always -1 spread). In practice, raising RR reduces the number of trades needed, cutting the total spread paid, which works in your favor. But raising RR also lowers the win rate, so at firms with a consistency rule, profit gets concentrated into fewer days, which works against you instead. This has the same structure as an earlier verification.
Q. If bigger risk is always better, is an all-in single shot optimal?
Looking purely at expected value, yes — but it doesn't work in practice. There's a daily loss cap that a single loss must not touch (at a 5% daily limit, 4% risk is near the ceiling), and there are minimum trading-day requirements too.
And when actually measured, 4% risk punched through the daily DD on 4 days. A gap jumping the stop-loss can occasionally cost 4x the intended amount in one shot, and the higher the risk, the more that single hit becomes fatal. Resolution does get faster (4 days per run), but risk around 2% turned out to be the realistic sweet spot.
Q. Is there any point to capping it at once a day?
No. Through Section 9 of this article, that's how it was set, but that's an experimental condition meant to isolate the per-trade economics — not a recommendation. Trading up to just short of the full daily loss allowance shrinks resolution from 139 days to 7-33 days, without lowering the pass rate. Leaving the allowance unused had zero upside — it was purely throwing away time and coupons.
Q. I have multiple accounts. Is it fine to use the same setting on all of them?
Better not to. Even with the same expected value, if the correlation between accounts is close to 1, adding more accounts doesn't cut variance by a factor of 12 — it stays piled up at 12x. In a one-year simulation across 12 accounts, the probability of finishing positive was 98% for "different seed per account" vs. 56% for "same setting on every account." On top of that, identically configured multiple accounts run straight into copy-trade detection, with the risk that all accounts' payouts get denied at once. See Section 12 for details.
Q. Wouldn't this violate the terms of service?
It doesn't fall under Fintokei's explicit bans (arbitrage, copy trading, HFT, gap trading). That said, every firm has a catch-all clause letting them deny payouts on the grounds of "lack of genuine trading," and there's no guarantee random entries won't be flagged as a problem. This article analyzes the structure of the expected value — it is not a recommendation to actually run this.
Sources
- Track360 – Prop Trading Industry Statistics 2026 (5-14% pass rates, 45% of funded accounts reach a payout)
- QuantVPS – Prop Firm Statistics 2026 (7% payout-reach rate at TopStep)
- FTMO – Forbidden Trading Practices
- Fintokei official site 🎁 (banned practices, plan terms, pricing)
- Each firm's rules and prices are from this site's comparison page database (as of August 2026)
※ Each firm's rules are revised frequently. Always check the latest version on the official site before purchasing.
Related articles
- What percentage of Phase 1 would a coin toss pass? The exact-solution answer comes out above 50%
- The consistency rule was targeting trend-following all along: testing 10,000 scenarios with a real EA
- Drawdown types explained: the difference between static DD, trailing DD, and daily DD
- High-variance strategies are actually well-suited to prop firms [Analysis]
- The complete prop firm guide: a roadmap from choosing a firm to getting paid
Written by
Hosono P | the prop firm strategist
I buy challenges with my own money and record everything through to the payout. Recorded payouts: ¥6.1M in total from Fintokei, Fundora and Funded7, plus $4,776 from The5ers (as of September 2026). Author of the semi-discretionary EA "ELDRA".