Mithqal — مثقال ← Back to the site
The measurement log

Six rounds. Most of them killed our own ideas.

Every trading product tells you what works. This page tells you what we tried, what it measured, and which of our own ideas died as a result. The rule we run under: an idea is adopted only if the measurement says so, and the measurement is published either way.

What one round actually is

A round takes one named idea — "trade only when the 4-hour trend agrees", "move the stop to breakeven at +1R" — and runs it through the same backtest the live engine is measured by: real gold bars, each style on its native timeframe, costs of $0.30 spread and $0.10 slippage per side charged on every fill, stop-before-target on any bar that spans both.

The number that decides is expectancy in R: the average result per trade, measured in multiples of the risk taken. And the interval around it is computed by resampling whole DAYS, not individual trades — because two trades taken an hour apart are not two independent experiments, and pretending otherwise is how 23 days of gold turn into a confident-looking verdict.

Round 0 — the gate: "give us more calls"

The most requested change we have ever received. We swept the trend gate from ADX 25 down to 22 and 20, on all three styles.

Verdict: the gate stays at 25. Loosening it to 22 buys the 5-minute style 30 more trades and destroys 88% of its edge (+0.037R → +0.004R per trade). More calls measured as worse calls. The patience counter in the app is the honest answer to that request; a looser gate is not.

Rounds 1–2 — entry filters: both hypotheses dead

H1, "only take a trade when the 4-hour trend agrees": hurts every style, on 5 years and again on 9 (5m: +0.017 → −0.027; 1h: −0.021 → −0.043; 4h: +0.147 → +0.098). The counter-trend entries we would have filtered out were carrying value. The mechanism story was simply wrong.

H2, "stand aside when the dollar is moving against the trade" (measured against the Fed's own broad dollar index): worse across the board. Combining H1 and H2 is worse still.

Verdict: entry filtering is exhausted. Two plausible, widely-repeated ideas, both dead on our own data.

Round 3 — the stop floor: a cost drag, named

If the 1-hour style bleeds, maybe its stops are too tight to survive the spread. We swept a minimum stop distance as a multiple of round-trip friction.

At 60× friction (a $30 minimum stop) the 1-hour bleed shrinks from −0.0182R to −0.0014R per trade — the loss was largely cost drag on thin stops. Neutralised. Not rescued: break-even is not an edge, and we do not sell break-even.

Round 4 — "the target just needs more time": dead

We multiplied every holding limit by two, four and eight. Expectancy did not move for any style, because timeouts are 0–2% of all exits. Trades resolve or die well inside their window; the ones that go nowhere were never going anywhere.

Round 5 — our own coach, on trial (and convicted)

The app has a trade coach. It used to advise moving the stop to breakeven at +1R and locking +1R at +2R — standard, comfortable, universally repeated advice. Nobody had ever measured whether OBEYING it beats holding the original bracket. So we did, over nine years and 3,394 fills.

5-minute: +0.0099R → +0.0032R (edge cut 67%). 1-hour: −0.0182R → −0.0296R (bleed worsened). 4-hour: +0.0837R → +0.0548R (edge cut 35%).

Breakeven amputates winners: trades that touch +1R and pull back to entry frequently went on to reach the target. A higher apparent safety, bought with a lower expectancy — the classic trap, and our own product was teaching it.

Action taken: that advice was downgraded from an instruction to a consideration, and it now states its own measured cost in the app, in the same sentence that offers it. Psychology has a price; the honest thing is to print the price.

Round 6 — the last lever: WHEN. Also dead.

If what you trade and how long you hold it cannot be improved, perhaps when you trade can. We swept every session combination a live filter could express — Asia, London, the London/New York overlap, New York, and every subset of them — for all three styles, over three years, each on its native timeframe.

Not one subset produced a decisive positive result for any style. The only intervals that excluded zero excluded it on the losing side. The most tempting number in the whole program was the 1-hour style trading the overlap alone: +0.100R per trade — with an interval of [−0.027, +0.227] on 371 trades. Adopting that would be fitting a window, not finding an edge, and we would rather print it here than sell it.

What this leaves standing — and it is not an edge

The same round measured our best number over a longer window, and it did not survive. The 5-minute style reads +0.037R over 180 days and −0.0015R over three years, with an interval of [−0.034, +0.029] across 1,900 trades. Two windows of the same bars disagree. That makes the good number a property of its window, not of the engine.

So the honest statement, the one this whole page exists to be able to make: no style in this product has a demonstrated edge. Not the 5-minute, not the 1-hour, not the 4-hour. Six rounds, every named lever we had, and that is the result.

What we do have is a workflow: a deterministic engine that publishes before the move and cannot edit afterwards, a record nobody can quietly amend, alerts that arrive when the engine speaks, a journal that grades your obedience to it, and a coach that states the measured cost of its own advice. That is what is for sale here. An edge is not, because we have not measured one.

This rule is enforced in our build, not in our conscience: a test fails the release if any sales sentence claims an edge while the measurement file says no style has one. The file and the test are both in the repository.

What is still open

Round 7 goes after our own assumptions rather than the market: we assume $0.30 of spread and $0.10 of slippage per side. If real fills are worse than that, every number above gets worse, and we would rather find that out than have a customer find it. After that: replacing single-window backtests with a rolling walk-forward, because the window sensitivity above is the argument against ever trusting one window again.

Every month, reconciled in public

The rounds above are backtests: claims about the future made by looking at the past. The live ledger is the only out-of-sample test that exists, and it accumulates whether anyone looks at it or not — so once a month it is looked at, in public, and compared against what we said to expect.

A month is not a verdict, and the pages say so in the same sentence as their numbers: below thirty resolved calls a month is published for transparency and explicitly refuses to be read as evidence in either direction. What actually watches for decay is the control chart on the live record, which halts publishing on its own if the rolling average breaks its limit.

Monthly reconciliations live at /measurement/<year-month>, one page per month, generated from the ledger.

The rule

Adopt only what the measurement blesses. Publish the result either way. Never let a number appear in our marketing that does not resolve to a query anyone can run against the public ledger.