“We have 4K cameras” tells you nothing about whether you can identify the person who walked out with the laptop. Resolution is a specification. What decides the outcome is how much of that resolution landed on the face — and that is a number you can calculate before anyone buys anything.
The formula
Pixels per foot is the camera’s horizontal pixel count divided by the width of the scene, in feet, at the distance where the image has to be useful.
a 1920×1080 camera across a 24 ft scene: 1920 ÷ 24 = 80 ppf
the same camera across a 96 ft yard: 1920 ÷ 96 = 20 ppf
a 3840-pixel-wide camera across that same yard: 3840 ÷ 96 = 40 ppf
Three things fall straight out of that line. First, the number depends on where you measure — a single camera delivers high density close in and low density far away, so a specification that does not name the plane has not specified anything. Second, doubling the scene width halves the density, which means moving the camera or changing the lens is usually cheaper than buying a bigger sensor. Third, vertical resolution barely features; the industry works horizontally because that is how coverage is described.
You will also see the same idea written in pixels per metre, which is how the international standards express it. Divide by 3.28 to get feet: 125 px/m is about 38 ppf.
How much is enough
The international video surveillance standard IEC 62676-4 defines operational requirement classes — monitor, detect, observe, recognise and identify — each with a pixel density target, rising in roughly a 1 : 2 : 5 : 10 : 20 progression from about 12.5 px/m to about 250 px/m. In feet, those land near:
- Monitor — about 4 ppf. Is there a crowd, and is it moving?
- Detect — about 8 ppf. Something is there, and it is probably a person.
- Observe — about 19 ppf. What is happening, and roughly what the person is wearing.
- Recognise — about 38 ppf. That is the day-shift supervisor, whom I already know.
- Identify — about 76 ppf. An image good enough to identify a stranger.
That is where the common North American shorthand comes from: roughly 40 to recognise, roughly 80 to identify. Treat these as design targets, not guarantees. They are the minimum pixel density at which the task is achievable under favourable conditions — not a promise that any scene meeting the number will produce a usable image.
The practical consequence is that identification is a location, not a system-wide setting. You do not identify people across a building; you identify people at the three or four points where they must pass, and you accept observe- or recognise-level coverage everywhere else. A design that claims identification everywhere has either not been calculated or has been priced for a different building.
Working a real opening
Take a main entrance. The objective is to identify people entering — not staff, whom you would recognise, but anyone. Work backwards.
- Start at the target, not the camera. The plane that matters is where a face crosses the doorway, roughly five to six feet above the floor.
- Pick the scene width you need at that plane. For identification, something in the region of four to six feet — one or two people abreast, not the whole lobby.
- Do the arithmetic. 76 ppf across 6 ft needs about 456 horizontal pixels. That is well inside any modern sensor — which tells you the constraint here is not resolution at all. It is the lens, the aiming and the light.
- Choose the lens to produce that scene width at that distance. Longer lens, narrower scene, higher density. This is the step people skip by mounting a wide-angle camera above the door and hoping.
- Check the angle. A camera high above the doorway pointing steeply down produces excellent images of the tops of heads. Mount lower, or set the camera back and shoot along the direction of travel.
- Then add the second camera — wide, observe-level — for context, because your identification camera sees almost nothing else.
Notice what did not happen: nobody specified a megapixel count. The identification job at a doorway is solved by geometry and light. Megapixels start to matter when you try to hold density across a wide scene — 38 ppf across a 50 ft room needs about 1,900 horizontal pixels, which puts a 1080p camera exactly at its limit with nothing left over.
Five things that ruin the number
Pixel density is necessary and not sufficient. Five factors regularly turn a correctly calculated camera into unusable footage:
- Motion and shutter speed. A walking subject under a slow shutter is a smear, whatever the pixel count. Freezing motion needs a faster shutter, which needs more light.
- Backlight. A glass entrance behind the subject silhouettes them. Wide dynamic range helps; aiming the camera somewhere else, or fixing the lighting, helps more.
- Compression. Aggressive bitrate caps and smart codecs preserve the static background and discard exactly the moving detail you needed. Check what the recorded file looks like, not the live view.
- Lens quality and focus. A sensor can only record what the optics deliver. A soft or mis-focused lens loses detail that no resolution recovers — and autofocus drifts over years.
- Infrared at night. Monochrome, a hotspot near the lens, reflective clothing blowing out, and effective range shorter than the datasheet. Site lighting outperforms it almost every time.
How to write it into a specification
Write the objective, the plane and the density — not the product. Something on these lines:
“Camera C-101 shall achieve not less than 76 pixels per foot of horizontal resolution across the target plane located at the north entry threshold, measured at 5 ft 6 in above finished floor, under the design lighting condition. The contractor shall submit a field-of-view calculation demonstrating compliance prior to ordering, and shall demonstrate the achieved image during acceptance testing.”
Three things that language does. It makes the requirement testable rather than aspirational. It puts the calculation on the party proposing the product, which is where it belongs. And it makes substitution safe: any camera and lens combination that meets the density is acceptable, so you are buying a result rather than a part number. That last property is worth more than it looks when a long lead time forces a change two months into the job.
If you want to see the trade-offs interactively, the camera selection decision tree walks the same logic by use case, and IP video 101 covers the frame rate, bitrate and retention side of the same decision.
Who wrote this, and what we sell. LA CCTV Supply provides security consulting and system design, sells the equipment and trains your people. We are not an installing contractor: installation is performed by your licensed contractor, except for small non-permitted work under $1,000 all-inclusive, which we can handle directly.
Nothing above is legal, code or accreditation advice. Requirements for your project are set by your Authority Having Jurisdiction, your engineer of record, your contracting officer or your Accrediting Official — and where a figure depends on your building, we have said so rather than inventing one.
