engineer, engineering, electrical engineer, electrical engineering, science, research, vehicle, cars, automobile, computer, laptop, software, testing, engineer, electrical engineering, software, software, software, software, software
Photo by This_is_Engineering on Pixabay

Stack Planning

Experimentation platforms

Choose an experimentation platform by checking delivery, assignment, measurement, privacy and the data your team needs.

An experimentation platform assigns eligible people to different experiences and connects those assignments to outcomes. Choose one by starting with the change your team wants to test: a marketing page, a product feature or a process spanning both. The platform needs a workable way to deliver the change, record exposure and measure an outcome the team can explain.

Start with a proposed experiment

Before comparing feature lists, write down the eligible audience, what receives a variation, the change each group sees, the primary outcome and any guardrail measure. Identify who can stop the experiment and who will interpret the result.

A landing-page test might compare two form layouts and measure completed enquiries. A feature test might compare onboarding flows for signed-in users and measure completion while monitoring errors. The questions are similar, but the work falls to different editors, developers and data owners.

DecisionWhat to ask in a demonstration
DeliveryCan the team make and preview this specific change?
AssignmentWho is eligible, what receives a variation, and how consistent is assignment?
MeasurementWhat records exposure, and how is the outcome connected to it?
ControlWho can pause a variation, and what will affected users see?
AnalysisCan the team inspect group sizes, metric definitions and uncertainty?

Match the platform to the work

Optimizely Web Experimentation is a candidate for website page tests. Its documented workflow includes an experimentation editor, URL or saved-page targeting, a site snippet, metrics and testing the experiment. Ask the site team to check snippet placement, event setup and how the variation behaves on the actual page.

LaunchDarkly Experimentation is a candidate when a change is delivered through a feature flag or configuration. Its documented setup connects variations and metrics to an experiment; for LaunchDarkly-hosted metrics, flag evaluation events establish exposure. A developer should show the actual evaluation and metric event. If independent event analysis is required, confirm the proposed export route: Data Export is an add-on for select plans.

Statsig documents exposure and custom events for computing metrics and generating experiment results. Its raw-event documentation describes ingesting events through its SDKs or HTTP API, data connectors or warehouse imports. Ask which route the team would use and who maintains the identifiers and metric definitions.

Optimizely's documented sequence requires at least one metric and its one-line JavaScript snippet in place before an experiment can run. Audiences decide who sees the test and the default audience is Everyone. Page triggers and conditions can control when a page activates, which suits single-page applications.

LaunchDarkly requires a flag or config with its variations, a metric, and a rule chosen when the experiment is built. The flag need not be toggled on to create the experiment but must be on before an iteration starts. The context kind the rule targets should match the experiment's randomisation unit.

Documented limits apply to that model. An experiment cannot run on a flag or config whose rules include an active guarded rollout, an active progressive rollout, a running Data Export experiment or a running warehouse native metrics experiment. It also cannot run on a migration flag. Multiple experiments can share a flag, with only one running per rule.

Statsig requires at least one unit identifier on every raw event, both to keep a user's experience consistent when they are allocated to control or test groups and to join exposure events with that user's custom events. Its default unit identifiers are User ID and Stable ID.

Platform comparison: Optimizely, LaunchDarkly, and Statsig

Best for
Website page tests
Delivery method
Site snippet + experimentation editor
Exposure tracking
JavaScript snippet events
Primary use case
A/B testing on web pages
Best for
Feature flag-based changes
Delivery method
Feature flag or configuration
Exposure tracking
Flag evaluation events
Primary use case
Testing feature rollouts in apps
Best for
Custom event-driven experiments
Delivery method
SDKs or HTTP API, data connectors, warehouse imports
Exposure tracking
Raw events with unit identifiers (User ID, Stable ID)
Primary use case
Advanced analytics and custom metric integration

Pros and cons of key experimentation platforms

  • Optimizely Web ExperimentationPro: Simple setup for website A/B tests Con: Limited to web; relies on JavaScript snippet placement
  • LaunchDarkly ExperimentationPro: Integrates with feature flags and app logic Con: Requires developer involvement; Data Export add-on needed for full visibility
  • StatsigPro: Flexible event ingestion and custom metrics Con: Requires robust event schema and identifier management

Check the report and data path

Trace a sample assignment through exposure, outcome, report and any required export. If an outcome arrives later from a CRM or warehouse, establish how it joins to the original assignment. A summary export and event-level data support different kinds of analysis.

For an Australian deployment, identify personal information collected or sent to a vendor. Have the organisation's privacy owner assess applicable notice, use, disclosure and overseas-recipient obligations. The treatment of a testing cookie depends on the actual data flow and circumstances; a vendor's consent option does not settle the legal question.

How results are reported

Statsig documents two measures in its experiment results: a p-value, where a value below 0.05 typically indicates statistical significance, and a confidence interval, where the effect is statistically significant if the interval does not overlap zero. Ask a candidate to show both on a past test.

The randomisation unit is the entity — a user, device or session — that is assigned to a control or test group, and every value of that unit must always receive the same variant. A unit that does not match where users are in the product when assignment happens causes crossovers between groups and skews results.

Key metrics and statistical reporting standards

Statistical significance threshold
p-value < 0.05
Confidence interval interpretation
Effect is significant if interval does not overlap zero
Randomisation unit
User, device, or session — must remain consistent
Australian privacy consideration
Personal information collected must comply with APPs under OAIC guidelines

Decide with one representative test

Run one test from the workflow you expect to use: configure a variation, confirm assignment and exposure, verify the primary metric in the report, and check any required export. Choose a platform only if the responsible team can complete that path.

Steps to run a representative test

  1. Configure a variationSet up the change using the platform’s editor or code interface
  2. Confirm assignmentVerify eligible users are consistently assigned to control or test groups
  3. Verify exposureCheck that events are recorded when users interact with the variation
  4. Check primary metric in reportEnsure the outcome is accurately tracked and displayed
  5. Validate required exportConfirm data export to CRM, warehouse, or analytics tool works as expected

Pre-decision validation checklist

  • Can the team deliver and preview the change?Yes / No
  • Is user assignment consistent and traceable?Yes / No
  • Are exposure and outcome events properly linked?Yes / No
  • Can the team pause the experiment safely?Yes / No
  • Is the primary metric visible and interpretable?Yes / No
  • Does data export meet compliance needs?Yes / No

In this guide

  1. Comparing feature testing with marketing page testingCompare how feature and marketing page tests are delivered, assigned, measured and controlled before choosing a workflow.
  2. Checking traffic allocation and reporting in an experiment platformAudit eligible traffic, variation splits, exposures, metric definitions and live changes before acting on an experiment report.
  3. Reviewing privacy and consent controls for testing softwareCheck cookies, identifiers, event flows, refusal and withdrawal behaviour when assessing testing software for an Australian team.
  4. Assessing whether a testing tool supports the team's data needsCheck identifiers, exposure and outcome events, reports, exports and data ownership before choosing a testing tool.

More from Stack Planning