Building My Own Tool Scanner — Part 4: Let There Be Light
By this point, I had something I was pretty happy with.
The camera was calibrated.
I could take a position in the image and convert it into a real-world position in millimetres.
The calibration had been tested rather than simply assumed to work.
So surely the difficult bit was over?
Not quite.
Because knowing the exact position of a pixel isn't particularly useful if you've detected the wrong bloody pixel.
The scanner needed to see the tool
The next job was segmentation.
In simple terms:
Which pixels belong to the tool, and which belong to the background?
That sounds fairly straightforward.
Put a tool on a contrasting surface, take a photograph and detect its outline.
And with some objects, it was.
Then I put shiny tools under the camera.
Ratchets, sockets, extension bars, chrome spanners...
Basically all the things normally found in a toolbox.
And shiny metal is remarkably good at making computer vision considerably more annoying than it needs to be.
Reflections are not your friend
A polished tool doesn't have one consistent appearance.
One part might look almost black.
Another part of exactly the same tool might be bright white because it's reflecting a light source.
Curved surfaces create highlights.
Edges pick up reflections from the surrounding environment.
The tool can even reflect the scanning surface itself.
To a person looking at the image, it's obviously still one object.
To software trying to decide whether a particular pixel belongs to the tool?
Not necessarily.
I could keep making the segmentation algorithm increasingly complicated...
Or I could make the image easier to segment in the first place.
Stop trying to photograph the tool
That eventually led to a fairly important change in how I thought about the scanner.
I don't actually care what the tool looks like.
I don't need to see the chrome finish.
I don't care about the manufacturer's logo.
I don't need to identify whether it's a ratchet or a pair of pliers.
For the scanner's purpose, most of that information is completely useless.
What I actually want is:
The silhouette.
So rather than trying to illuminate the tool nicely, I started concentrating on illuminating the background.
Put a bright, controlled surface underneath the tool and suddenly the problem changes.
Instead of:
“Which collection of reflections belongs to this shiny object?”
I'm asking:
“Where does this bright background stop?”
That's a considerably easier question.
Backlighting the scanning area
The scanning surface effectively became a diffuser.
Light is introduced underneath it, producing a bright background while the tool blocks that light.
From the camera's point of view, I now have a dark silhouette sitting against a much brighter surface.
That makes thresholding and mask generation considerably more reliable.
The scanner doesn't need to understand the tool.
It just needs to find the interruption in the light.
Simple.
Mostly.
Because naturally, solving one problem created another. 😂
Uniformity matters
If one part of the scanning area is significantly brighter than another, the software isn't seeing one consistent background.
It's seeing a gradient.
That can affect where the threshold falls and, ultimately, where the detected edge of the tool ends up.
And remember what I'm eventually doing with that edge:
Manufacturing something from it.
A few pixels of segmentation error can become a dimensional error in the resulting pocket.
So illumination wasn't just about making the image look nicer.
It became part of the measurement system.
That meant experimenting with light position, diffusion and the scanning surface itself to make the background as even as reasonably possible.
Hardware can fix software problems
This became another useful lesson from the project.
My first instinct was naturally:
Improve the image processing.
More filtering.
More clever thresholding.
More code.
But every bit of software I added was compensating for poor input data.
Improving the physical lighting attacked the problem before the image ever reached Python.
Cleaner input.
Simpler segmentation.
More reliable edge detection.
Less code trying to guess what the camera should have seen.
Sometimes the best way to improve an algorithm is to give it better data.
From image to mask
With the lighting under control, the scanner could finally produce the thing I actually needed:
A clean binary mask.
Background on one side.
Tool on the other.
That mask became the bridge between the physical object sitting on the scanner and everything that came afterwards.
Calibration told me where the edge was.
Illumination allowed me to reliably find the edge.
Now I needed to turn that edge into something useful.
Because ultimately I still had the same goal I'd started with:
I wanted a Gridfinity insert without manually modelling the bloody tool.
And having a nice mask on a computer screen wasn't going to organise my toolbox.
Next: Part 5 — Okay, I've Scanned It. Now What?
The scanner could finally capture useful geometry.
The next problem was everything that happened after pressing Scan.
That started with a fairly simple web-based builder.
It did not remain fairly simple for long. 😂
