Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Real-World Mathematics: Photography, Optics, Perspective, Resolution and Image Geometry

Application of Mathematics in Real-World Usage · Guide 18 · BTT Mathematics Hub

A camera turns three-dimensional space into a two-dimensional image. That process is not a perfect copy. Distances from the lens become image distances, object size becomes image size, perspective shrinks distant objects, a sensor samples the projected image into pixels, and a display or print maps those pixels onto another physical surface. Photography is therefore an unusually clear example of representation Mathematics.

This guide develops a simplified thin-lens and pinhole-style geometric model. Every camera, sensor, focal length, subject size and pixel count below is a fictional or generic teaching example unless a physical equation is linked to a source. Real photographic systems include lens thickness, distortion, focus tolerances, diffraction, sensor response, processing and many other effects. The purpose is to understand the geometry, not to calibrate equipment or prescribe exposure settings.

Keep four spaces distinct: world space contains the subject, optical space contains object and image distances, sensor space measures millimetres across the projected image, and pixel space counts discrete samples. Confusing two spaces often produces a calculation with the right arithmetic and the wrong meaning.

The thin-lens equation connects object distance, image distance and focal length

OpenStax gives the thin-lens relation 1/f = 1/dₒ + 1/dᵢ, with sign conventions depending on the lens and image type. In a simple positive-distance teaching case with a converging lens, take focal length f = 50 mm and object distance dₒ = 1,000 mm.

Then 1/dᵢ = 1/50 − 1/1000 = 19/1000 mm⁻¹, so dᵢ = 1000/19 ≈ 52.63 mm. The image plane lies slightly beyond the focal length because the object is at a finite distance.

If the object moves very far away, 1/dₒ approaches zero and dᵢ approaches f. This limiting behaviour is a useful check on the equation. A result far smaller than f for this positive real-image model would deserve investigation.

Magnification is a ratio of image size to object size

For the same thin-lens model, OpenStax writes magnification as m = hᵢ/hₒ = −dᵢ/dₒ. With dᵢ≈52.63 mm and dₒ=1,000 mm, m≈−0.05263.

A 200 mm tall object would therefore form an image about 10.53 mm tall in magnitude. The negative sign indicates inversion under the stated sign convention; the size ratio is about 5.263%.

Magnification is dimensionless because the length units cancel. Using metres for the object and millimetres for the image without converting would corrupt the ratio by a factor of 1,000.

Perspective can be modelled with similar triangles

In a pinhole-style projection, image size is approximately proportional to focal distance and object size, and inversely proportional to subject distance. A simple relation is y = fY/Z, where Y is object height, Z is distance from the projection centre and y is image height on the sensor.

Suppose Y=2.0 m, Z=10 m and f=50 mm. Convert Y and Z to the same unit or use their ratio 2/10=0.2. The image height is 50×0.2=10 mm.

Move the same subject to 20 m while holding everything else fixed. Image height becomes 5 mm. Doubling distance halves projected linear size in this simple model.

Perspective change is not the same as crop

Suppose a subject remains at 10 m and produces a 10 mm image height. Cropping the final image to use only the central half of the sensor can make the subject occupy a larger fraction of the displayed frame, but the projected geometry on the sensor was still created from the original camera position.

Moving the camera closer changes ratios between objects at different depths. Cropping does not. This distinction matters when interpreting “same framing.” Two pictures can show a subject at the same displayed size while carrying different perspective relationships.

Mathematically, perspective depends on the camera-to-object geometry; crop is a later selection inside the image plane. They are different transformations.

Field of view is an angle from sensor size and focal length

For a rectilinear pinhole model, angular field of view across sensor dimension s can be written θ = 2 arctan[s/(2f)]. This is a geometric model, not a claim that every real lens maps angle perfectly this way.

Using a 36 mm sensor width and 50 mm focal length gives horizontal field of view 2 arctan(36/100)≈39.60°. Using a 24 mm height gives about 26.99° vertically.

Changing sensor dimension while keeping focal length fixed changes the captured angle. Changing focal length while keeping sensor fixed also changes it. Field of view is therefore a relationship between both quantities.

Doubling focal length does not halve the angle exactly

Because field of view uses an arctangent, it is not exactly inversely proportional to focal length. With a 36 mm sensor width, 50 mm gives about 39.60°. At 100 mm, the field is 2 arctan(36/200)≈20.41°, slightly more than half of 39.60°.

For narrow angles, inverse approximations can be useful. For wider angles, the nonlinearity becomes more noticeable. This is a good example of when a simple proportional intuition is close but not exact.

A graph of field of view against focal length would fall rapidly at small focal lengths and flatten as focal length grows. The shape comes from the arctangent function.

Aspect ratio is shape, not resolution

An image 6000×4000 pixels has aspect ratio 6000:4000 = 3:2. An image 3000×2000 also has ratio 3:2 but one quarter as many pixels.

Aspect ratio describes proportional width and height. Pixel count describes the number of discrete samples. Two images can share shape while differing greatly in resolution.

Cropping a 6000×4000 image to a square 4000×4000 changes aspect ratio from 3:2 to 1:1 and discards 8 million pixels. It does not magically change the physical characteristics of the lens or sensor that made the original image.

Megapixels come from two dimensions multiplied

A 6000×4000 image contains 24,000,000 pixels, or 24 megapixels when “mega” is used as one million. A 3000×2000 version contains 6 million pixels.

Halving both width and height produces one quarter the total pixel count. This square-law scaling is the same Mathematics that appears in map area and image enlargement.

Doubling megapixels does not mean doubling each linear dimension. To double total pixels while preserving aspect ratio, multiply each dimension by √2.

Pixel pitch converts sensor millimetres into sampling density

Suppose a 36 mm sensor width is sampled by 6000 pixels. The average horizontal pixel pitch in this idealised rectangular model is 36/6000 = 0.006 mm = 6 micrometres.

If another sensor has the same 36 mm width but 9000 pixels across, the horizontal pitch becomes 4 micrometres. More samples fit into the same physical distance.

Pixel pitch alone does not determine image quality. It is one geometric sampling quantity. Optics, noise, processing, diffraction and many other factors are outside this simple calculation.

Print size converts pixels into samples per physical inch

If a 6000-pixel image width is printed at a chosen 300 pixels per inch, print width is 6000/300 = 20 inches. At 200 pixels per inch, the same image produces 30 inches.

The image file has not gained pixels. The physical spacing assigned to them has changed. A higher pixels-per-inch value produces a smaller print for the same pixel dimensions.

Whether a particular sampling density is sufficient for a real viewing condition is a perceptual and production question. The arithmetic only maps digital samples onto a physical size.

Sensor crop factor is a ratio of diagonals

Consider a 36×24 mm reference sensor. Its diagonal is √(36²+24²)≈43.27 mm. A geometrically similar 24×16 mm sensor has diagonal √(24²+16²)≈28.84 mm.

The ratio 43.27/28.84 = 1.5. Under the same lens and camera position, the smaller sensor captures a smaller central portion of the projected image. The 1.5 factor describes the diagonal size relationship.

Multiplying focal length by 1.5 can be useful for comparing approximate fields of view between these two sensor formats, but it does not physically transform the lens into one with a different focal length.

Similar framing from different distances changes depth relationships

Place subject A at 5 m and background B at 10 m. In the simple pinhole model, their image scales are proportional to 1/5 and 1/10, so A appears twice the linear scale of an equal-size B.

Move the camera back so A is at 10 m and B at 15 m, then use a longer focal length to restore A’s displayed size. Their relative scale is now proportional to 1/10 and 1/15, so A is only 1.5 times the scale of equal-size B.

The longer focal length restored subject framing, but the changed camera position altered the depth ratios. This is why “zooming” and “moving” are not geometrically equivalent transformations.

The inverse-square law is a different geometry from perspective

For an ideal point source spreading uniformly through three-dimensional space, intensity scales as 1/r². Doubling distance reduces the modelled intensity to one quarter.

Perspective image height, by contrast, scales approximately as 1/r in the pinhole model. Doubling subject distance halves linear image height, not quarters it.

Both relationships involve distance, but the exponents differ because one concerns linear projection and the other concerns distribution over spherical area. Similar-looking “distance effects” can have different mathematics.

Aperture number is itself a ratio

In the simplified f-number definition N=f/D, focal length f is divided by effective aperture diameter D. A 50 mm focal length at f/4 corresponds to an idealised aperture diameter 50/4 = 12.5 mm.

At f/5.6 the corresponding diameter is about 8.93 mm. Aperture area is proportional to diameter squared, so the f/4 opening has about (5.6/4)²=1.96 times the area of f/5.6, close to a factor of two.

This is geometry only. Actual transmission also depends on lens design and losses. The familiar sequence of f-numbers reflects square-root-of-two scaling because area depends on the square of diameter.

Exposure time scales linearly with time in a fixed-rate model

If all other factors and scene brightness are assumed fixed, an idealised sensor receiving a constant light rate for 1/125 s collects twice the time-integrated quantity of one exposed for 1/250 s.

The times are 0.008 s and 0.004 s. Their ratio is 2:1. This is a simple accumulation model.

Real exposure systems can introduce sensor response, shutter behaviour and other effects. The arithmetic is intended only to show the direct relationship between constant rate and duration.

Combining aperture and time requires multiplying factors

Suppose a classroom model changes aperture area by a factor of one half and exposure time by a factor of two. Their product is one, so the modelled integrated quantity stays constant.

Changing each control by “one stop” in opposite directions is therefore represented by reciprocal powers of two. The compensation is multiplicative rather than additive.

This does not claim two real photographs will look identical. Motion blur, depth of field, diffraction, noise and other properties can change even when one aggregate exposure quantity is held similar.

Resolution limits can be stated as bounds rather than guarantees

Suppose a target detail projects to 2.4 pixels across under a simplified sampling model. If a classification rule requires at least 3 pixels across, the target fails the rule. Increasing image scale by 25% gives exactly 3 pixels.

The threshold is a decision rule, not a universal law of visibility. Real detectability depends on contrast, optics, noise and processing.

Threshold Mathematics is still useful because it makes the requirement auditable. A vague claim such as “high enough resolution” becomes a testable condition once the chosen metric and threshold are declared.

Cropping changes available output size

A 6000×4000 image cropped to 4500×3000 retains 13.5 million pixels. Relative to the original 24 million, the retained fraction is 13.5/24 = 56.25%.

At 300 pixels per inch, the cropped width supports 15 inches rather than 20 inches at the same sampling density. Enlarging both to the same physical width would require the cropped image to use fewer pixels per inch.

This shows why crop, output size and sampling density form a linked triangle. Changing one while holding another fixed forces the third to change.

Rotation changes bounding-box dimensions

Take a rectangle of width w and height h rotated by angle θ around its centre. The axis-aligned bounding box has width |w cosθ|+|h sinθ| and height |w sinθ|+|h cosθ|.

For w=6000, h=4000 and θ=90°, the bounding box becomes 4000×6000. At 45°, the bounding width and height are both approximately (6000+4000)/√2≈7071 pixels.

The image content itself has not grown merely because the bounding box is larger. Empty corner regions may appear after rotation unless the image is cropped or extended.

Homogeneous coordinates make perspective projection algebraic

Advanced computer vision often represents a 3D point (X,Y,Z) projectively. In a simple normalized pinhole model, image coordinates are x=X/Z and y=Y/Z before focal and sensor scaling.

Multiplying X, Y and Z by the same nonzero constant leaves x and y unchanged. The projected location depends on ratios, not the absolute coordinate scale.

This invariance explains why perspective geometry is naturally connected to ratio and linear algebra. It also shows the singularity at Z=0: division fails because the point lies on the projection plane through the camera centre in this idealised coordinate model.

A complete image-geometry report keeps the transformation chain visible

For a defensible calculation, state object size and distance, focal-length or projection assumptions, sensor dimensions, pixel dimensions and output dimensions separately. Then identify where cropping, scaling or rounding occurs.

A 10 mm projected subject, for example, can become 1667 pixels tall on a sensor sampled at 6000 pixels across 36 mm if the horizontal and vertical sampling pitch is assumed equal at 6 μm. That pixel count can then be mapped to a display size independently.

The same final displayed size can come from very different optical and digital chains. Mathematics makes the chain explicit instead of treating “zoom” as one undifferentiated operation.

Practice: twenty photography Mathematics questions

  1. For f=50 mm and object distance 1000 mm, find image distance using the thin-lens equation.
  2. Using question 1, find magnification.
  3. A 200 mm object has magnification magnitude 0.05. Find image height.
  4. Using y=fY/Z, find projected image height for a 1.8 m subject at 9 m with f=50 mm.
  5. What happens to projected linear size if subject distance doubles with everything else fixed?
  6. Find horizontal field of view for a 36 mm sensor width and f=50 mm using 2 arctan[s/(2f)].
  7. Find pixel count for 6000×4000.
  8. Reduce the aspect ratio 6000:4000.
  9. If a 36 mm width contains 6000 pixels, find pixel pitch in micrometres.
  10. How wide is 6000 pixels at 300 pixels per inch?
  11. Find the diagonal of a 36×24 mm rectangle.
  12. Find the crop-factor ratio between diagonals of 36×24 mm and 24×16 mm.
  13. At f=50 mm and f-number 4, find idealised aperture diameter.
  14. Compare ideal aperture areas at f/4 and f/5.6 through their diameter-square ratio.
  15. Compare exposure times 1/125 s and 1/250 s.
  16. A crop keeps 4500×3000 from 6000×4000. What percentage of pixels remains?
  17. At 300 pixels per inch, find output width of the 4500-pixel crop.
  18. Under normalized perspective, compare image coordinate x=X/Z for (X,Z)=(2,10) and (4,20).
  19. A target detail covers 2.4 pixels and a rule requires 3. By what scale factor must image size increase?
  20. Why is moving closer to a subject not geometrically identical to cropping or using a longer focal length from the original position?

Worked answers

  1. About 52.63 mm. 1/dᵢ=1/50−1/1000.
  2. About −0.05263. Use −dᵢ/dₒ.
  3. 10 mm.
  4. 10 mm. The object-to-distance ratio is 1.8/9=0.2.
  5. It halves. In the simple projection model image size is inversely proportional to distance.
  6. About 39.60°.
  7. 24 million pixels.
  8. 3:2.
  9. 6 μm.
  10. 20 inches.
  11. About 43.27 mm.
  12. 1.5.
  13. 12.5 mm.
  14. f/4 has about 1.96 times the ideal aperture area of f/5.6.
  15. 1/125 s is twice as long.
  16. 56.25%. 13.5 million divided by 24 million.
  17. 15 inches.
  18. Both give 0.2. Scaling X and Z by the same factor preserves the ratio.
  19. 1.25. 3/2.4.
  20. Because camera position changes the depth ratios between objects. Cropping changes only the selected image region, and changing focal length changes projection scale from the same viewpoint.

Sources and connected applications

For the thin-lens equation, image distance and magnification conventions, see OpenStax: Thin Lenses. The sensor dimensions, pixel counts, aperture examples, crop factors and projection cases here are original teaching constructions built from standard geometric relationships rather than claims about a specific camera model.

Continue with Music, Rhythm, Frequency, Ratios and Sound; Retail, Pricing, Discounts, Margins and Break-Even; and Agriculture, Crop Yield, Irrigation, Sampling and Resource Planning. Return to the BTT Mathematics Hub.