RoboQuant
RoboQuant engine

Backtesting

Roboquant runs the compiled .rqc artifact for your strategy against historical CME data. The Backtest tab combines execution controls, a live progress stream, chart replay, strategy logs, trades, and performance analysis in one view.

How a run works

.rq source → Compile → .rqc artifact
                         ↓
Backtest request → isolated worker pod
                         ↓
Historical bars, ticks, and optional L2 depth
                         ↓
Compiled engine → orders, fills, equity, logs, drawings
                         ↓
Stored result → chart, Replay, metrics, trades, analysis

The artifact is loaded before any data is processed. Parameters are injected when the run starts, so changing a parameter does not require recompilation.

Starting a backtest

From Strategies → Backtest:

  1. Select the compiled entry artifact.
  2. Choose the symbol, timeframe, start date, and end date.
  3. Choose a fill model.
  4. Set starting capital, commission, slippage, and leverage.
  5. Review the parameters generated from #[param(...)] fields.
  6. Click Run for the fastest result or Replay to watch the strategy progress across the chart.

Every saved run records the selected entry file, parameters, date range, data source, L2 setting, capital, commissions, slippage, and leverage. Reopening a run restores that configuration.

Multiple entry files

A strategy can hold several .rq files. Backtest, Optimize and Deploy each ask which entry file to run when there is more than one; the choice is saved with the run, so reopening it runs the same file. If a saved run's entry file has been deleted or renamed, choose an available one before running again.

Fill models

Fills settingMarket timingData usedBest for
OHLCVOrders execute with bar-level simulationHistorical barsFast research and hourly/daily systems
TickTriggers and fills follow the historical trade tapeIndividual trade printsIntrabar entries, SL/TP, and trailing-stop behavior
Order Book (L2) — barMarket orders walk the depth snapshot at the bar eventBars + MBP-10 depthBar strategies where liquidity and order size matter
Order Book (L2) — tickTick timing plus depth-walked market fillsTrade tape + MBP-10 depthHighest-fidelity short-window execution research

Tick and L2 data are available only for supported CME symbols and covered dates. L2 coverage is a recent rolling window and is intended for short, usually single-day studies. The engine reports average levels walked and slippage cost when book fills are used.

Limit and stop orders use their trigger semantics. Market orders walk visible depth when L2 is enabled; if the visible ladder cannot fill the requested size, the run uses the configured fallback slippage model.

In OHLCV mode a stop-loss that the bar gaps through fills at the bar's open. A take-profit fills at its price and never better, but never outside the bar's range: a take-profit already through the market exits at the bar's low (longs) or high (shorts).

Bar, tick, and timer behavior

Strategy hookRequired run dataBehavior
on_barAny fill modelRuns at each completed bar
on_tickTick or L2 tickRuns for each trade print; the strategy must return true from wants_ticks
on_timerTick or L2 tickRuns on the simulated timer grid set in on_init
on_tradeAny fill modelRuns for order, fill, position, and close transactions

In tick mode, ctx.bar(0) is the forming bar as known at that tick. Indicators still read the last completed bar. This prevents final-bar values from leaking into earlier tick decisions.

Run vs Replay

Both buttons execute the same engine and produce the same final result.

  • Run computes without pacing the chart and is best when you only need the result.
  • Replay streams bars, intrabar tick substeps, indicators, trades, stop/target rails, trailing-stop movement, drawings, equity, and logs while the run computes.

Replay speed is a display control; it does not change execution results. A tick replay uses its actual intrabar event timestamps, so entries and exits appear at their real point inside the candle instead of being snapped to the candle boundary.

Results

The summary includes core measures such as:

CategoryExamples
ReturnNet P&L, total return, monthly returns
RiskMax drawdown, Sharpe, Sortino
Trade qualityWin rate, profit factor, average trade, largest win/loss
ExcursionMAE and MFE where available
Statistical contextAlpha, beta, and probabilistic Sharpe where available
ExecutionTrade list, commissions, slippage, L2 levels walked
DiagnosticsStrategy logs, drawings summary, engine errors

The result chart includes entries and exits, equity and drawdown curves, and strategy drawings. Multi-symbol results retain the executing symbol on every trade and store candles per leg.

The runner also records run statistics computed from the full bar series (not the downsampled chart): total, warm-up and active bars, first and last bar time (bar open, UTC), and bars per ET hour. Together with entry hours and order-rejection counts they form a price-free diagnostics block that explains zero- or low-trade runs. Multi-symbol runs count the primary leg's bars; results from older runners report the statistics as unavailable.

Metrics describe a historical simulation, not a forecast. Always compare them with costs, trade count, exposure, regime concentration, and out-of-sample behavior.

Costs and contracts

The engine applies the instrument contract specification automatically. Order size is a contract count; P&L uses the configured point multiplier.

Commission is charged on each side:

commission = flat per trade
           + per-contract charge × contracts
           + percent charge × price × contracts × multiplier

Use per-contract commission for futures. Percentage commission is primarily for legacy spot-style runs. Slippage is applied according to the selected fill model; L2 market fills use the real visible ladder before falling back to the configured percentage.

Multi-symbol backtests

Compiled MultiStrategy artifacts run across a list of symbols with one merged on_bars timeline.

  • The first symbol is the primary chart symbol.
  • Every leg needs data for the requested timeframe and date range.
  • Orders and account reads are addressed by symbol name.
  • Results include per-leg trades and OHLC series.
  • Multi-symbol execution is OHLCV bar mode only; tick and L2 modes are rejected.

See Strategies → Multi-symbol strategies for the authoring pattern.

Optimization

The Optimize tab uses the same artifact, data, costs, and execution engine as a normal backtest. Parameters come from the artifact's #[param(...)] schema.

Search methods

OptimizerUse it for
Grid SearchExhaustive, reproducible sweeps over a manageable parameter space
Bayesian (TPE)Efficient search when the full grid is large
Genetic (CMA-ES)Continuous or irregular search spaces
Multi-Objective (Pareto / NSGA-II)Trade-offs such as Sharpe versus drawdown instead of one winning score

Validation methods

ValidationWhat it does
SingleScores every trial over the full selected history
IS / OOSOptimizes on the first portion and validates on the untouched remainder
Walk-ForwardRepeats rolling in-sample optimization and out-of-sample validation

Market universes

The market selector can build a symbol × timeframe matrix for any search and validation method. Every parameter candidate is evaluated independently on every selected market case, then receives one aggregate score. The same parameters are shared across the whole matrix; this searches for a robust configuration rather than a different optimum per symbol.

The Optimize tab uses a dispersion-adjusted objective by default:

maximize metric: score = weighted mean − 0.25 × standard deviation
minimize metric: score = weighted mean + 0.25 × standard deviation

Results retain each market's metrics alongside the aggregate mean, dispersion, and worst case. A failed market invalidates the candidate by default. The planned evaluation count is the number of parameter candidates multiplied by the number of symbol/timeframe cases, so widening either axis increases runtime and is included in the 10,000-evaluation safety limit. One optimization can contain at most 40 market cases.

Market-universe optimization is distinct from a multi-symbol portfolio strategy: a market universe runs the ordinary strategy independently on several markets, while a MultiStrategy trades several symbols together in one account. The two modes cannot be combined.

Tick fills support the same symbol × timeframe matrix as OHLCV: every case gets its own tick tape, built once per distinct symbol and shared across cases on that symbol. Tick runs over a matrix of two or more markets are capped at 3,650 symbol-days in total (distinct symbols × days in the date range, about 10 symbol-years); a single-market tick run keeps the usual date-range limit, so a wide matrix over a long window may need a narrower date range or fewer symbols.

Market-only mode

Clearing every parameter range while keeping two or more market cases selected switches the run to market-only mode: the strategy runs once, at its default or fixed parameter values, on each selected market, and the markets are ranked against each other by the target metric instead of ranking parameter sets. This is useful for finding which symbols and timeframes a strategy already suits, before tuning parameters on the best ones.

Market-only mode always uses grid search over a single candidate and the "Single (full history)" validation — there is nothing to sample or fit with no parameter range enabled, so Bayesian/Genetic/Pareto search and IS/OOS/Walk-Forward validation are unavailable in this mode. Unlike a parameter sweep, a market that fails to produce a result is dropped from the ranking instead of invalidating the run; the run only fails if none of the selected markets produce a result.

Pareto search currently uses Single validation. Multi-symbol optimization also uses bar data; uneven multi-leg walk-forward windows are not supported. Tick optimization works for a single market or a symbol/timeframe matrix (not for on_bars multi-symbol portfolio strategies) and can be substantially slower than OHLCV because every trial replays the tape.

Optimizer results include best parameters, full-grid or trial distributions, parameter heatmaps/surfaces, and walk-forward or Pareto views where applicable. A promising optimum should be re-run as a normal backtest with the exact selected configuration before deployment.

Data universe

The dashboard exposes Roboquant's licensed CME futures universe: equity-index, interest-rate, FX, livestock, and crypto futures roots. CME is the only market-data source; availability still varies by symbol, timeframe, and data type.

Historical lookup follows this order inside the platform:

  1. Local Parquet cache in the worker.
  2. Resampling from cached minute data when possible.
  3. Roboquant's S3 Parquet store.
  4. The configured upstream provider for environments that enable it.

If a date range is outside stored coverage, the run fails with a data-availability error instead of silently changing the dates.

Contract rolls on continuous series

Continuous CME series (for example NQ rather than a dated contract) switch from the expiring contract to the next one at 00:00 UTC on the roll date published for that root. The raw prices jump by the gap between the two contracts at that moment. The compiled engine does not treat that jump as a market move:

  • At the roll, open positions, stop-loss and take-profit levels, trailing levels and resting limit/stop orders are all shifted by the published gap, so no order fills and no P&L is booked on the jump.
  • A trade held across the roll shows its entry at the new-contract-equivalent price.
  • The result carries an account warning for each roll applied during the run. A run that crosses more than three rolls gets one summary warning instead; the full list is in the result's market data (market_data.rolls).

Dated contracts do not roll and are unaffected.

Known limits:

  • Bars that span a roll (for example 1W or 1M) apply it at the next bar in the bar lane, so an order can still fill on the gap inside that bar.
  • Python (.py) strategies run on the interpreted engine for bar and tick backtests (Run button, MCP and AI chat). It only warns about the rolls in the run; it does not shift positions or orders, so the jump can fill orders and move open P&L. Compiled .rq strategies are covered as described above.
  • A strategy that re-places orders from its own stored absolute price levels puts them back at old-contract levels after a roll. Recompute levels from current prices or ctx state instead.

Reproducibility checklist

Before comparing two runs, confirm that both use the same:

  • compiled artifact version;
  • symbol list, timeframe, and date range;
  • fill model and L2 setting;
  • capital, leverage, commission, and slippage;
  • strategy parameters;
  • validation split, when optimizing.