We Ran a General-Purpose AI Against a Model Trained on Real Faults. The Accuracy Gap Was Enormous.
Everyone wants to know if an off-the-shelf AI can just "look at a photo" and flag the defect. We ran the head-to-head on real field-inspection images. Here are the numbers, and what they mean for your inspection costs.
The question every ops team is asking
If you run field inspections, you have probably been asked some version of this in the last year: can't we just point a general-purpose AI at a photo and have it tell us what's wrong? The pitch is seductive. No training, no setup, no specialist team. Upload an image, get an answer.
It's a fair question, and it deserves a fair test. So we ran one. We took a general-purpose AI vision tool, the kind that can describe almost any image in plain language, and put it head-to-head against a model trained specifically on the faults our customers actually inspect for. Same images. Same defects. Same task. We wanted to know, honestly, whether the convenient option was good enough.
How we tested it fairly
We used a real defect-inspection dataset drawn from fiber and telecommunications field inspection, the kind of imagery a technician captures on a phone in the field. We asked both approaches to do the same job: find the specific defect types in each photo and pinpoint where they were, not just say something looked off somewhere in the frame.
On one side, off-the-shelf cloud vision AI working from a general understanding of the world. On the other, a purpose-trained computer-vision model that had learned from examples of the exact faults in question. Both were scored the same way, on how precisely they identified real defects and how many real defects they caught.
The accuracy result
The gap was not close. The purpose-trained model hit 95.6% precision and 95.6% recall: when it flagged a defect, it was almost always a real one, and it found almost all of the defects that were there.
The best run from the general-purpose LLM vision tool landed at roughly 3-5% precision and 3-5% recall on the same task. In practical terms, the general model identified only a small fraction of the real faults, and most of what it did flag wasn't useful. For an inspection workflow, that's the difference between a tool you can rely on and one you'd have to double-check by hand on every single photo.
Purpose-trained: 95.6% precision and recall. General-purpose vision: roughly 3-5% on the same images.
Why the general model struggled
This isn't because general-purpose AI is bad. It's because it's general. These models are built to describe anything, which means they have never specifically learned what your particular faults look like, how they differ from normal wear, or exactly where in a busy photo they tend to appear.
So when asked to locate a defect, the general model tended to gesture at vague regions of the image rather than pinpoint the actual fault. That's fine if you want a rough description. It's not fine if a technician needs to know precisely which component failed and where, so the right repair gets dispatched the first time.
The cost twist most people miss
Accuracy is only half the story. The economics matter just as much, and they cut the same way.
General-purpose cloud vision is typically billed by the image, and often by how much the model has to analyze within each one. As an illustration, on typical token pricing that runs roughly EUR 25-40 per 1,000 images, and the cost climbs with every additional photo and every detection. That's manageable for a pilot. It becomes a real line item when you scale.
A self-hosted, purpose-trained model flips that math. Once it's built, it runs at a flat, predictable rate, with effectively no incremental token cost per image. You're not paying more each time your team takes another photo.
What this means at scale
Picture thousands of inspections a week, which is routine for a busy telecommunications or utilities operation. At that volume, per-image billing keeps growing in lockstep with your activity, while the accuracy problem multiplies: a model that misses most real faults isn't just expensive, it lets defects through.
A purpose-trained model gives you the opposite profile. High accuracy that holds up across volume, and a cost that stays flat whether you run a hundred inspections or a hundred thousand. The more you inspect, the wider the gap in your favor.
How Forge gets you to the high-accuracy side
Here's the part that usually surprises people: building the accurate model doesn't require a team of data scientists. That's exactly what Forge by SkillsBase is for.
Forge lets your own team, the people who already know your faults, train a custom computer-vision model on your real defects and deploy it to technicians' phones in days, not months. No code, no consultants, no ML PhDs. Just your inspection knowledge, turned into a model that actually works in the field.
If you're weighing an off-the-shelf AI assistant against a model trained on your own faults, the numbers above are the reason this choice matters. Book a demo and we'll show you what a purpose-trained model looks like on your inspections.
How we measured this. These figures were measured on a focused, real-world fault dataset for a specific set of defect types, not a universal benchmark. Results will vary by use case, image quality, and the defects you inspect for. The cost figure is illustrative, based on typical token pricing rather than a guaranteed rate.