Service

Data Annotation & AI Training Data

Trained annotation teams for images, video, 3D LiDAR and point clouds, computer vision, robotics, text, audio, documents and LLM training data, run with written guidelines, gold sets and a separate review pass.

We give you a trained team not just a labelling tool

People who already do this every day Apex Automation Team runs trained annotation teams across every common modality. You tell us the model you are training and what it has to recognise, and we assemble the team, write the labelling guidelines with you, run a calibrated pilot batch and then scale. Production and review are done by different people, so nothing reaches you having been checked only by the person who produced it.

We take any size of engagement A single batch of a few thousand frames, a standing team that clears a weekly queue, or a named group that works only on your project and learns your edge cases. We supply the people, the supervision and the reporting, so you are managing one delivery lead rather than a roster of annotators.

Image annotation

  • Bounding boxes and rotated or oriented boxes
  • Polygons and polylines, including multi-part shapes and holes
  • Semantic segmentation, instance segmentation and panoptic segmentation
  • Amodal segmentation for occluded objects, plus matting and trimaps for hair, glass and smoke
  • Keypoints, facial landmarks and skeletal pose with visibility and occlusion flags
  • Cuboids on 2D images where depth and yaw matter
  • Single-label, multi-label and hierarchical classification
  • Per-object attributes such as colour, material, state, truncation, occlusion level and damage severity
  • Scene graphs and relationship triples
  • Image quality grading for blur, exposure, noise, glare and focus
  • Captioning, dense captioning, visual question answer pairs and referring expressions for vision language models
  • Text in image work: detection, transcription, reading order and per-region script identification

Video annotation

  • Frame by frame labelling and keyframe annotation with interpolation
  • Single and multi-object tracking with persistent track identities
  • Segmentation with tracking, so masks carry an identity across the clip
  • Occlusion and re-entry handling, drift correction and identity switch prevention
  • Re-identification across cameras, scenes and time gaps
  • Action recognition, temporal action localisation and phase segmentation
  • Spatio-temporal action detection where the label follows the box through the clip
  • Gaze, head pose, blink and distraction labelling for driver and operator monitoring
  • Shot boundary, scene change and highlight detection

3D LiDAR and point clouds

  • 3D cuboids with position, dimensions and orientation, tracked across sequential frames
  • Point level semantic and instance segmentation, and 3D panoptic segmentation
  • Lane centrelines, dividers, curbs, guardrails and road boundaries as 3D polylines
  • Ground plane and drivable free space labelling
  • 2D to 3D linking so the same object carries one identity in camera and cloud
  • Multi-sensor fusion across LiDAR, camera, radar and inertial data
  • Calibration checks, time synchronisation verification and ego-motion annotation
  • Multi-sweep aggregation, registration and sequence stitching
  • HD map layers: lane graphs, signs, signals, stop lines, crosswalks and road furniture
  • Static and dynamic point separation with motion state attributes

Robotics and embodied AI

  • Teleoperated demonstration collection and human demonstration capture
  • Egocentric and wrist camera video collection across randomised scenes and tasks
  • Episode and trajectory labelling with clean task boundaries
  • Skill and sub-action segmentation down to reach, grasp, lift, transport, insert, place and release
  • Grasp annotation: grasp points, 6-DoF grasp poses, gripper width and approach vectors
  • Object pose annotation and pose tracking, hand keypoints and bi-manual coordination
  • Affordance labelling and part level affordance masks
  • Articulation annotation for doors, drawers, handles and joint ranges
  • Language instruction grounding, so natural instructions line up with the trajectory
  • Success, partial success and failure labelling with a failure mode taxonomy
  • Deliberate failure and recovery episodes, which is the data most collections are missing
  • Safety and constraint violation labelling for collisions, drops and excessive force

Text audio and documents

  • Named entity recognition including nested and overlapping spans, entity linking and relation extraction
  • Intent and utterance labelling, slot filling, dialogue acts and out of scope handling
  • Sentiment at document, sentence and aspect level, emotion, toxicity and abuse labelling
  • Classification, topic labelling, coreference, part of speech and dependency parsing
  • Search relevance grading and question answer pair creation
  • Summarisation quality review for faithfulness, coverage and hallucination
  • Translation and localisation quality with span level error typing and severity
  • Verbatim and clean transcription, timestamping and forced alignment
  • Speaker diarisation, overlapping speech, voice activity detection and speaker metadata
  • Phonetic transcription, prosody, emotion and paralinguistic tagging
  • Audio event detection with onsets and offsets, wake word and near miss samples
  • Document work: field and key value extraction, table structure, layout regions, reading order, handwriting transcription and document classification
  • Personal data identification, redaction and de-identification

LLM and agent training data

  • Instruction and response pair authoring, including domain expert answers and reasoning traces
  • Multi-turn dialogue authoring, clarification turns and refusal writing
  • Tool use and function calling traces with schema valid arguments
  • Grounded answer authoring with citations tied to the supporting passage
  • Pairwise preference ranking, best of n ordering and reward model data
  • Rubric based scoring and per-dimension grading for helpfulness, correctness, safety, tone and formatting
  • Critique and revise pairs, where the reviewer writes both the fault and the fix
  • Model output evaluation, side by side comparison and regression checks between versions
  • Hallucination and citation verification at the level of individual claims
  • Adversarial prompting and safety probing against a written policy, with severity grading
  • Agent trajectory annotation: step correctness, tool choice, argument correctness, loops, recovery and hallucinated success
  • Browser and computer use demonstration capture with the actions and targets, not only the video
  • Multilingual prompt collection and in-language evaluation by native speakers

Work that needs subject knowledge

Where the label needs someone who knows the subject Medical imaging with organ and lesion contouring, measurements and report labelling. Geospatial and drone imagery with building footprints, land cover, crop boundaries, infrastructure inspection and change detection. Agriculture with crop and weed separation, growth stage, disease severity and counting. Retail with shelf and product recognition, planogram compliance and out of stock detection. Manufacturing with surface defect segmentation, severity grading and assembly verification. Sports with player tracking, jersey recognition, pose and event labelling. Security with intrusion, loitering, abandoned object and threat detection. Insurance with damage segmentation, severity grading and repair or replace calls.

How we hold the quality

Guidelines are written before volume starts Nothing scales until the schema is settled. We write the label taxonomy, the attribute list, the decision tree for the awkward cases and a do and do not gallery, then keep an edge case register so the twentieth annotator answers a hard case the same way the first one did. Guidelines are versioned, so any batch can be traced to the rules that were in force when it was produced.

Nobody starts on live data Annotators sit a qualification test scored against gold answers before they touch your project. A paid pilot batch runs first and is measured rather than delivered, which is where guideline holes get found cheaply.

Production and review are done by different people Work passes from annotator to reviewer to lead audit. Gold items are seeded into the live queue so quality is measured continuously rather than at the end. Overlap sampling lets us report real agreement figures instead of asserting accuracy.

Reported in numbers you can check Agreement measured with the statistic that suits the label type, spatial accuracy by intersection over union or Dice, tracking accuracy by identity metrics, and defects grouped by cause rather than counted as one flat rate. Disagreements go to an adjudicator, and repeated defects are traced back to the guideline, the training or the individual.

Delivered in your format COCO, YOLO, Pascal VOC, KITTI, nuScenes, OpenLABEL, CVAT, Label Studio, DICOM and segmentation formats, GeoJSON and shapefiles, CoNLL and JSONL, or a schema you define. We agree the exact output record before work starts so the data lands ready for training rather than needing conversion.

Confidentiality is standard on every project Signed agreements at company and individual level, least privilege access, named reviewers, and restricted environments where the data requires it.

Frequently asked

What kinds of data do you annotate?

Images, video, 3D LiDAR and point clouds, text, audio, documents, and the human data used to train and evaluate language models and agents. That covers boxes, polygons, segmentation, keypoints and pose, tracking and re-identification, 3D cuboids and sensor fusion, robotics demonstrations and trajectories, entity and intent labelling, transcription and diarisation, field and table extraction, preference ranking, rubric grading, evaluation and adversarial testing.

How do you keep quality consistent across a large team?

The schema is settled before volume starts, with written guidelines, a decision tree for hard cases and an edge case register that grows as questions are answered. Annotators pass a qualification test scored against gold answers before touching your data, a pilot batch is measured rather than delivered, and production and review are always done by different people. Gold items sit inside the live queue so quality is measured continuously, and overlap sampling gives real agreement figures rather than a claimed accuracy number.

Can you work on a project that needs subject knowledge?

Yes. Medical imaging, legal, financial, engineering, scientific and code work are staffed with people who have the relevant background, and review on those projects is done by someone qualified to disagree with the annotator. We tell you who is doing the work and what they are qualified in.

What format do you deliver in?

Whatever your training pipeline expects. COCO, YOLO, Pascal VOC, KITTI, nuScenes, OpenLABEL, CVAT and Label Studio exports, DICOM and segmentation formats, GeoJSON and shapefiles, CoNLL and JSONL, or a custom schema. We agree the exact output record before work starts so the data arrives ready to train on.

Do you handle confidential or sensitive data?

Yes, under signed agreements at both company and individual level, with least privilege access, named reviewers and restricted working environments where the data requires it. Personal data can be identified and redacted as part of the work rather than afterwards.

All Apex Automation Team FAQs →