TradinSolutionsEXECUTION LAYER
Start Free Trial →
HOME/BLOG

Notes from building execution infrastructure.

Prop-firm drawdown arithmetic, broker symbol suffixes, contract rolls, and what actually breaks when you copy a trade across five platforms.

← ALL POSTS
Platform Tutorials28 Sept 2026 · 8 min · TradinSolutions

A Demo Testing Protocol for Automation: The Two-Week Checklist

Fourteen days, a fixed list of things to check on each of them, and a written pass or fail at the end. The protocol itself, not the argument for running one.

Day nine of a demo test is where the useful failures live. The first week is configuration — you are still finding the setting you typed wrong. The second week is when the thing runs unattended long enough to hit a rollover, a widened spread, a symbol you never tested, and a restart you did not plan. Most traders stop testing on day four, because by day four it looks like it works.

This post is the protocol, not the argument for having one. It assumes you have already decided to test before going live and want a concrete list to work through. It applies to any piece of automation attached to a MetaTrader, cTrader or similar terminal: a copier, a risk manager, a journal sync, an execution bridge.

Before day one: fix the test conditions

A test where you change the configuration halfway measures nothing. Freeze the following and write them down.

  • The exact build. Version number, file hash if you have one, and the date you installed it. If the vendor ships an update mid-test, the test restarts.
  • The full settings block. Every input, copied verbatim into your notes. Not "default risk" — the actual number.
  • The symbol list. The instruments you will genuinely trade, not a convenient subset. Gold and indices behave nothing like EURUSD under stress, and they are where sizing bugs surface.
  • The host. Same VPS, same region, same terminal build you intend to use live. A test on your laptop tells you about your laptop.
  • The pass conditions, written before you start. Deciding what counts as a pass after you have seen the results is not a test, it is a rationalisation.

TIP

Take a screenshot of the settings dialog on day zero. When something behaves oddly on day eleven, the first question is always "did I change something" and the screenshot answers it in two seconds.

The daily loop: eight checks, five minutes

Run these every trading day of the test. The point is not any single check — it is that fourteen consecutive days of the same eight checks make a drift visible that a one-off glance never would.

  1. 01Is it still running? Terminal connected, automation enabled, no dialog waiting for a click. Record uptime, not just "yes".
  2. 02Did every position get a stop? Scan open and closed positions for a missing or zero stop level. One naked position in fourteen days is a fail, not a curiosity.
  3. 03Does the size match the rule? Take one trade from the day, compute by hand what the lot size should have been from equity and stop distance, and compare. Do this on a different symbol each day.
  4. 04Any rejected or requoted orders? Read the journal and experts tabs, not the trade list. Rejections often leave no trace on the chart.
  5. 05Any orphans? A position open on one account with no counterpart where there should be one, or a pending order that should have been cancelled and was not.
  6. 06What was the worst fill? Record the largest gap between intended and actual entry price for the day, in pips and in currency.
  7. 07Log size and error lines. Grep the day's log for the word "error" or its equivalent. Zero is the expected answer; anything else gets written down with the timestamp.
  8. 08Equity curve sanity. Not "is it up" — is the shape consistent with the number and size of trades you saw. A jump you cannot account for is the single most important thing you will find all fortnight.

Each check takes seconds. The discipline is doing all eight on the day you are busy.

The event calendar: what the fortnight must contain

Fourteen days of quiet markets is not a test of anything. Deliberately schedule the window so it includes at least three of the following, and note the date of each.

  • A high-impact release. CPI, non-farm payrolls, or a central-bank rate decision. You are watching spread behaviour, slippage, and whether any news filter you configured actually fired.
  • A session rollover. The daily server rollover, where spreads widen and swap is applied. Hold at least one position through it on purpose.
  • A weekend. Carry one small position from Friday to Monday, once, so you see how the automation handles the gap and the reconnect.
  • A deliberate restart. Kill the terminal mid-session and bring it back. Nothing should replay, nothing should duplicate, and any open position should be recognised rather than re-opened.
  • A connection drop. Disable the network adapter for two minutes. Same expectation: reconnect, reconcile, do not duplicate.

The last two are the checks nobody runs and the ones that catch the expensive bugs. Duplicate execution after a reconnect is the classic failure in copying software, and it only ever appears when you interrupt something.

WARNING

Run the restart and disconnect tests on a demo account only, and only while you are watching. If a tool does duplicate on reconnect, you want to see it happen on a screen, not discover it in a statement.

What to record, and in what shape

A demo test produces a document, not a feeling. Keep one table, one row per day.

DayTradesMissing SLSize checkWorst slipErrorsNotes
140ok (XAUUSD)0.4 pip0baseline
260ok (EURUSD)1.1 pip0—
950ok (US500)6.8 pip2CPI 13:30, both errors "off quotes"

Two weeks of that table is worth more than any amount of remembering. It also gives you a baseline: when something odd happens in month three of live trading, the honest question is whether it is new, and the table is the only thing that can answer it.

The numbers above are an illustrative shape for the log, not benchmarks to aim at. Your worst acceptable slippage depends on your stop distance, and a 6.8 pip slip on a 15-pip stop is a different event from the same slip on a 90-pip stop.

Pass conditions

At the end of day fourteen you write one of two words. The conditions below are the minimum; add your own, but do not remove any.

  • Every position had a stop. No exceptions, no "that one was manual".
  • Every hand-computed size matched, within rounding to the broker's lot step.
  • Zero duplicates across the restart and disconnect tests.
  • Zero orphans — nothing left open or pending that should have been closed or cancelled.
  • Every error line is explained. Not "there were only two" — you know what caused each one and whether it can recur.
  • Slippage is inside the model you assumed, including on the high-impact day.
  • Uptime is what you expected. If the automation was down for six hours on day eight and you did not notice until day ten, the monitoring failed even if the software did not.

Anything short of that is a fail, and a fail means fix the cause and restart the fourteen days. Restarting is annoying, which is the entire reason the protocol exists: it is much less annoying than the alternative.

The tests that are worth running twice

Three checks are cheap and catch disproportionately expensive problems.

Symbol mapping, on every instrument. Place one minimum-size trade on each symbol in your list and confirm it lands on the instrument you meant. Brokers name the same market differently — a gold symbol with a suffix, an index that is DJ30 at one venue and US30 at another. A mapping that silently fails is worse than one that errors, because the chart you are watching looks perfectly healthy.

Sizing at the extremes. Force a very tight stop and a very wide one. Tight stops are where a sizing formula produces an absurdly large lot, and a maximum-lot clamp is the thing that should catch it. If there is no clamp, that is a finding.

Behaviour when a limit is hit. If you configured a daily-loss halt, make it fire. Reduce the threshold to something trivially small for one afternoon and confirm the halt actually stops new entries and that it releases at the time you expect. A guard nobody has ever seen trigger is an assumption, not a guard.

After the fortnight

A pass on demo is permission to move to a live account at minimum size, not permission to go to target size. Demo fills and live fills differ, and the gap between them is the thing the next stage measures. The staged approach — demo, then live small, then live at size, each with its own pass condition — is covered properly in demo-to-live-soak-test-protocol, and this fortnight is the first of those stages done in detail.

Where this fits

The staged framework this checklist plugs into is in demo-to-live-soak-test-protocol; read that for the gates either side of these fourteen days. If you are choosing a demo environment to run the test in, and want to know what a demo can and cannot tell you about trend tools and timing, cfd-demo-platforms-trend-tools covers the platform side.

NEXT