How much labelled data do we need?
For fine-tuning a pre-trained model, a few hundred labelled examples can produce a viable baseline. The amount depends on task complexity and class diversity. We assess your data in discovery before committing to a target.
Can you handle proprietary document formats?
Yes. We handle PDFs, scanned documents, Excel files, and custom formats. Layout-aware models work with multi-column and tabular structures that simpler extraction tools cannot.
What evaluation metrics do you use?
We use metrics appropriate to the task — mAP for detection, F1 for classification and NER, BLEU/ROUGE for generation. Critically, we agree on the eval metric and threshold before training begins, not after.
Do you provide the trained model weights and pipeline?
The engagement defines which model artifacts, training code, evaluation assets, and documentation are deliverables, subject to the licenses and restrictions of any third-party models or components.