Skip to main content
Validraft

Statistical validation

Trade Count and the Illusion of Independent Evidence

A long trade log may contain repeated exposure to the same days, the same market shock, and the same underlying bet.

4 min readResearch and simulation only
Sample sizeDependenceUncertainty
Dark storm clouds gathering over the New York City skyline.
Photo (cropped and colour-graded): Anthony Quintano · CC BY 2.0 · source

Working definition

Independent evidence concerns how much distinct information observations contribute to a stated estimate. Trade count measures recorded activity; it does not, by itself, establish the uncertainty around that estimate.

01

Start with what the rows share

Imagine a hypothetical strategy that opens positions in twenty technology stocks on the same morning and holds them for five sessions. A broad sector move can affect all twenty outcomes. If another cohort opens the next morning, much of its holding period overlaps the first. Forty trade records are useful for execution review, but treating them as forty independent repetitions ignores their shared exposure.

That example does not imply a universal formula for dividing the trade count. Dependence can vary with the statistic, portfolio construction, holding period, and market environment. Begin by showing common entry dates, simultaneous positions, repeated instruments, and concentration in sectors or event clusters. A reviewer should be able to see where apparently separate observations came from.

02

Choose the unit that matches the claim

A claim about implemented portfolio performance calls for a portfolio return series with contemporaneous positions, cash, and costs. A claim about an event response may require event-level outcomes and explicit handling of overlapping windows. A trade-level win rate answers a different question again. Moving between these units without explanation changes what is being estimated.

Aggregation is useful, but it is not a certificate of independence. Daily portfolio returns may remain serially dependent. Andrew Lo's work on Sharpe-ratio statistics shows why uncertainty and annualization depend on assumptions about the return process; the familiar square-root-of-time conversion cannot simply be applied to every correlated series.

  • State whether the estimate describes trades, events, sessions, or the combined portfolio.
  • Report calendar coverage and overlapping exposure alongside the raw observation count.
  • Keep the chosen return frequency consistent across estimates, benchmarks, and uncertainty calculations.

03

Resample a structure the strategy could have produced

An independent draw of individual trades can dismantle the time clusters that the study is meant to evaluate. A block bootstrap instead resamples stretches of a series to preserve some local dependence. The R boot package documents both fixed-length blocks and blocks with randomly varying lengths. Neither choice automatically captures every dependency or makes a changing market stationary.

For a strategy driven by common shocks, a practical review might compare uncertainty using chronological portfolio-return blocks and separately inspect outcomes by event cluster. Document the resampling unit, block-length rationale, random seed, and sensitivity to reasonable alternatives. These are study-design choices, not defaults to select according to whichever produces the narrowest interval.

04

Describe the limits of the available history

More resampling iterations can reduce numerical variability in the simulation, but they do not supply additional historical market episodes. If most of the apparent edge comes from a small set of overlapping positions during one period, that concentration should remain visible even when the backtest contains many rows.

We would want a delivery to pair the headline estimate with its uncertainty method, exposure concentration, calendar coverage, and relevant limitations. Where independence cannot be defended, avoid translating a large trade count into a confident verdict. The useful next step may be a longer untouched evaluation or a narrower research claim, rather than another decimal place in the performance table.

Practical takeaways

  • Inspect shared time windows and exposures before interpreting the number of trades.
  • Match the observation unit and uncertainty method to the claim being tested.
  • Treat block length and clustering choices as explicit assumptions with sensitivity checks.
  • Additional bootstrap paths do not create additional market history.

Want an independent read?

Test the claim, the data, and the implementation together.

Validraft scopes the hypothesis, checks feasibility, and delivers a descriptive validation report with visible evidence and limitations. Research and simulation only; never investment advice.