Service 02 / Synthetic data

Data made for
the job ahead.

Create datasets for testing, model training or gaps in coverage. We define what the data needs to represent, generate it to a specification and check it against the intended use.

Discuss a dataset

A dataset with
a defined purpose.

01

Structured records and time series

Create tabular records with defined fields, relationships and constraints. Useful for software testing, scenario exploration or augmentation when the approach is supported by downstream evaluation.

02

Text and labelled examples

Build task-specific text examples, classification labels, extraction pairs or evaluation cases. Review consistency, duplication, coverage and the quality of labels.

03

Images and controlled variations

Scope image generation, augmentation or rendered examples around the visual conditions that matter. Define how annotations and variations will be checked.

04

Rare cases and testing scenarios

Produce invalid inputs, boundary conditions or less common scenarios to exercise a system. Each test case should have a clear purpose and expected behaviour.

Rules first.
Rows second.

Illustrative retail test records. The useful part is the specification behind them: valid ranges, consistent relationships and intentional test cases.

Illustrative synthetic records · not customer data
record_idcategoryquantityunit_priceorder_totalscenario
SYN-001home2450.00900.00standard_order
SYN-002electronics11299.001299.00single_item
SYN-003office10025.002500.00bulk_order
SYN-004home0450.000.00zero_quantity_boundary

Example constraints: non-negative quantities, known categories, unique IDs and order_total = quantity × unit_price. Actual project rules are agreed in the dataset specification.

More rows are
not the whole goal.

01

Structure and consistency

Validate types, ranges, required fields, relationships and domain rules. Label invalid test records deliberately so they are not confused with training data.

02

Coverage and diversity

Check the combinations and cases specified in the brief. Inspect duplicates, class balance and unintended patterns introduced by the generator.

03

Usefulness for the intended task

Where applicable, compare downstream model performance or software test coverage with an appropriate baseline. A dataset that looks plausible may still be unhelpful.

04

Privacy and permitted use

Discuss source-data permissions and potential disclosure risks. Checks may include duplication and similarity review. Formal privacy guarantees require a specific method and agreed threat model; synthetic data alone provides no such guarantee.

The data.
And the recipe.

  • ✓
    A versioned dataset

    Agreed formats such as CSV, JSONL or image files with annotations.

  • ✓
    A generation workflow

    Scripts or documented configuration, where included in the scope.

  • ✓
    A validation report

    Checks performed, observed failures and unresolved limitations.

  • ✓
    A data card

    Schema, source assumptions, intended use and usage restrictions.

What does your dataset need to do?

  • Is it for model training, software testing or evaluation?
  • Which fields, labels, formats or conditions should it include?
  • What data or source material can be used with permission?
  • Which rare cases and coverage gaps matter most?
  • How will you judge whether the dataset is useful?

Some datasets can be generated from a schema and rules. Others need suitable source material or a simulator. We check feasibility before committing to the full dataset.

Tell us what’s missing.

Start with the use case, the data format and the cases you need to cover.

Start a project brief