AI Training Data: Collection, Annotation, and Post-Training
The model is the cheap part. A model is a few lines of configuration and a bill; the dataset it learns from is months of work by people who understand the task, and that is where most AI projects actually stall. We collect data to your specification, label it, and — because building AI systems is our main business — we are the people who later have to train on it. That changes how the specification gets written.
Why AI projects stall on data rather than models
Teams arrive with a model choice already made and a dataset problem they have not named yet. The pattern is consistent.
- The labelling guidelines were written before anyone had seen the hard cases, so the hard cases are labelled inconsistently.
- The collected data matches the demo environment rather than the deployment environment.
- Nobody measured agreement between annotators, so nobody knows whether the labels mean anything.
- The dataset covers the common case thoroughly and the failure case not at all, which is the reverse of what training needs.
- The format has to be rewritten before it can enter the training pipeline, and the rewrite loses information.
What we do
Four capabilities, sold separately or as one programme.
- Data collection to a written specification — video, image, audio, document, and scripted scenarios.
- Egocentric and exocentric capture for robotics and embodied AI, recorded as synchronised pairs.
- Annotation across image, video, 3D point cloud, document, and audio.
- LLM post-training data: supervised fine-tuning demonstrations, preference ranking, and evaluation.
The team that labels it also trains on it
Most annotation vendors hand over a dataset and never learn whether it worked. We build AI systems as our main business, which means a dataset we produce is one we may have to fine-tune on, evaluate, and defend in production. That is not a marketing distinction — it decides what goes into the labelling guidelines, because the person writing them knows which ambiguities will surface later as model failures.
How we keep quality measurable
Quality claims are worth nothing without a measurement behind them, so we instrument the work rather than asserting the outcome.
- Gold-standard items seeded into every batch, scored continuously rather than audited at the end.
- Multi-pass review, with the second pass blind to the first.
- Inter-annotator agreement reported as a number, not a claim.
- Automated pre-labelling where a model can do the first pass, with people correcting rather than starting from nothing.
- Guidelines versioned like code, because a guideline change silently invalidates everything labelled before it.
Security and access
What we can state plainly, we state; what we cannot, we leave out.
- Every person touching your data is covered by an NDA before they see any of it.
- Access is granted per project, not per employee, and revoked when the project ends.
- Client data stays in an environment separated from our other work.
- Data is deleted on completion, on a schedule agreed in the contract.
Frequently asked questions
- What is the smallest project you take?
- A pilot batch. It is the right first step regardless of eventual size, because it is the only way to find out whether the guidelines survive contact with real data.
- How does a pilot batch work?
- You send a representative sample and an acceptance criterion. We label it, report the disagreements we hit, and rewrite the guidelines against them. The disagreements are the point — a pilot that produces none was too easy to be informative.
- Who owns the data and the labels?
- You do, including anything derived from them. We retain nothing after the agreed deletion date, and we never reuse client data to build datasets of our own.
- What happens when annotators disagree?
- Disagreement is escalated, not averaged. A senior reviewer adjudicates, and the resolution goes back into the guidelines so the same case is not decided twice.
- Do you work in languages other than English and Vietnamese?
- Talk to us about the language and the volume. We would rather tell you we cannot staff it than staff it badly.
Start with a pilot batch
Send a sample and what you need it to become. We will tell you what the specification is missing.
Get in touch
