Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

The Mathematics of Bukit Timah Rainforest | Sampling → Biodiversity → Uncertainty → What the Data Can Actually Tell Us

Bukit Timah Nature Reserve contains more life than any visitor can count in one walk. NParks currently lists more than 1,240 plant species, over 140 bird species, over 60 butterfly species and over 30 mammal species in the Reserve. Another official account notes that this 163-hectare fragment — only about 0.23% of Singapore’s land area — is home to more than half of the country’s native plant species.

Those numbers are remarkable.

They also create a mathematical problem.

How do we know?

No scientist has stood at the entrance, pressed a button and made every plant, bird, beetle, mammal, fungus and microorganism in Bukit Timah line up to be counted. The forest is too complex, too dynamic and too imperfectly observable for that.

Instead, scientists sample.

They define plots. They walk transects. They repeat censuses. They measure tree trunks. They identify leaves. They record sightings. They compare years. They build databases. They estimate what the observed data can support — and, just as importantly, what it cannot.

That makes Bukit Timah rainforest a beautiful lesson in one of Mathematics’ most important responsibilities:

How do we reason about a whole system when we can only observe part of it?

This article belongs to the What about Bukit Timah? hub. The ecological owner remains How Bukit Timah Nature Reserve Works, while Alfred Russel Wallace in Bukit Timah owns the historical naturalist story. Here, our job is narrower: the Mathematics of observing biodiversity without pretending observation is perfect.


A Forest Is Not a List Waiting to Be Completed

Students often meet data as if the numbers already exist in a neat table.

There are 40 values. Find the mean.

There are 100 responses. Draw a pie chart.

There are 20 measurements. Calculate the standard deviation.

Real scientific data begin earlier than that.

Before the table exists, someone has to decide:

  • what should count as an observation;
  • where observations should be made;
  • how large a sample should be;
  • which times of day or seasons should be included;
  • how repeated observations will be handled;
  • what happens when a species is present but not detected;
  • how identification errors will be corrected;
  • what population the sample is supposed to represent;
  • how uncertainty will be reported.

That is why a biodiversity survey is not merely Biology with numbers attached.

It is a measurement system.

1. Population and Sample: The First Separation

Suppose our question is:

What tree species occur in Bukit Timah Nature Reserve?

The conceptual population might be all relevant trees in the Reserve under a specified definition.

But observing every tree repeatedly forever is not practical.

So researchers observe a subset: the sample.

This sounds simple until we notice that the sample is not automatically representative merely because it is large.

A thousand trees beside an easy trail may tell us less about the entire Reserve than a smaller but carefully distributed design covering different elevations, slopes, soils and forest types.

Sample size matters. Sampling design matters too.

2. Why Ecologists Use Plots

A forest plot creates a bounded observation window.

Inside that boundary, researchers can define consistent rules: which trees are included, what trunk diameter is measured, how individuals are tagged, how positions are recorded and when the next census occurs.

NParks’ current Tropical Forest Ecology Research Programme describes two 2-hectare Forest Dynamics Plots inside Bukit Timah Nature Reserve. Established in 1993 and 2004 respectively, they are re-censused regularly and form Singapore’s contribution to the global ForestGEO network of long-term forest plots.

Two hectares is 20,000 square metres.

A square of that area would have side length:

√20,000 ≈ 141.4 metres

Real research plots need not be perfect squares, but the calculation gives us scale. A 2-hectare permanent plot is large enough to contain substantial forest structure and small enough to survey systematically.

Yet even two 2-hectare plots together cover only 4 hectares of a 163-hectare reserve.

As a simple area fraction:

4 ÷ 163 ≈ 0.0245

about 2.45% of the Reserve’s area

That calculation does not mean the plots somehow “represent exactly 2.45% of biodiversity”. Area fraction and biodiversity representation are different quantities. It means only that the permanent plots occupy that approximate fraction of the Reserve’s stated area.

Again, the denominator has a job.

3. Why One Plot Is Not Enough for Every Question

Bukit Timah is heterogeneous.

Different parts of the Reserve differ in topography, disturbance history, forest type, moisture, edge exposure and vegetation structure. If all measurements came from one convenient patch, the sample could be precise about that patch and misleading about the larger landscape.

An NParks-linked plant-diversity study used 52 plots, each 20 metres by 5 metres, along nine transects covering a range of topography and forest types across Bukit Timah Nature Reserve.

Each small plot had area:

20 × 5 = 100 m²

Across 52 such plots, the nominal sampled area was:

52 × 100 = 5,200 m²

That is 0.52 hectares.

The number is useful, but the spatial design is at least as important. Spreading plots across different conditions allows researchers to compare forest types rather than confusing local conditions with reserve-wide patterns.

This is a crucial statistical lesson:

A sample can be large and biased. A smaller sample can be more informative if its design matches the question.

4. Random, Systematic and Stratified Sampling Are Different Tools

School statistics often introduces several sampling methods as vocabulary. Bukit Timah lets us see why the distinctions matter.

Random sampling

Locations are selected using a random mechanism so that selection is not consciously steered toward convenient or interesting places.

Randomness helps protect against certain biases, but random locations may still be difficult or unsafe to access, and a simple random design may undersample rare habitat types.

Systematic sampling

Measurements are taken at regular spatial intervals — for example, every fixed distance along a transect.

This can give good spatial coverage, but the design must be considered carefully if the landscape itself has repeating structure.

Stratified sampling

The study area is divided into meaningful groups — perhaps forest types, elevation bands or disturbance histories — and each group is sampled deliberately.

This is often powerful in heterogeneous landscapes because it prevents a dominant habitat from swallowing the sample while rare but important conditions disappear from view.

None of these methods is universally “best”.

The method has to fit the inference.

5. Species Richness Is Simple to Define and Hard to Measure Perfectly

One common biodiversity measure is species richness: the number of distinct species observed in a defined area or sample.

If a sample contains individuals from 27 distinct plant species, observed richness is 27.

The arithmetic is easy.

The inference is not.

Suppose there are rare species present in the forest but absent from our plots. Observed richness is then less than the true richness of the larger target area.

Suppose we sample twice as much area. We will often detect additional species simply because we looked harder and wider.

So when comparing two sites, raw richness can be unfair if sampling effort differs.

More observed species can mean more biodiversity. It can also mean more sampling.

Statistics exists partly to separate those possibilities.

6. Species Accumulation Curves: What Happens as We Keep Looking?

Imagine sampling one forest plot, then a second, then a third.

At first, many species may be new to the dataset. Later plots will contain a mixture of previously observed species and occasional new ones.

We can graph:

  • x-axis: sampling effort, number of plots or sampled area;
  • y-axis: cumulative number of species observed.

The resulting species accumulation curve often rises quickly and then begins flattening as new observations increasingly repeat species already found.

If the curve is still rising sharply, our inventory is clearly incomplete.

If it begins flattening, additional effort is producing fewer new species — though this still does not prove every species has been found.

This graph teaches a broader mathematical principle:

The value of the next observation depends on what the previous observations have already revealed.

7. Rare Species Are a Detection Problem

A common beginner mistake is:

“We did not see it, therefore it is not there.”

That conclusion is often too strong.

A species may be present but difficult to detect because it is rare, seasonal, nocturnal, high in the canopy, quiet, cryptic, underground, small or simply absent from the sampled locations during the observation window.

This introduces detection probability.

Suppose, hypothetically, a species is present at a site and a particular survey method has a 0.4 probability of detecting it on one independent visit.

The probability of missing it once is:

1 − 0.4 = 0.6

If the same detection probability applied independently across three visits, the probability of missing it all three times would be:

0.6³ = 0.216

So even after three visits, there would still be a 21.6% chance of failing to detect the species under those hypothetical assumptions.

The example is not a measured Bukit Timah detection probability. It illustrates the logic.

Absence of observation and evidence of absence are not automatically the same thing.

8. Independence Is an Assumption, Not a Magic Word

The calculation above used 0.6³.

That multiplication assumes the three detection events behave independently under the model.

But field observations may not be independent.

  • Weather can make all visits simultaneously less effective.
  • The same observer may repeat the same identification bias.
  • A species may move between visits.
  • Seasonality may alter detectability.
  • Nearby plots may share environmental conditions.

When dependence exists, multiplying probabilities as if observations were independent can create false confidence.

This is one of the most important habits in probability:

Before multiplying probabilities, ask why you are allowed to multiply them.

9. Richness and Abundance Are Different Questions

Imagine two hypothetical forest plots.

Plot A: 100 individuals belonging to 10 species, with each species represented by 10 individuals.

Plot B: 100 individuals belonging to 10 species, but one species contributes 91 individuals and each of the remaining nine species contributes one.

Both plots have species richness 10.

They do not have the same distribution of abundance.

This is why biodiversity cannot always be compressed into a single species-count number.

Depending on the research question, scientists may care about richness, evenness, relative abundance, rarity, functional groups, phylogenetic relationships or changes through time.

Different metrics answer different questions.

10. Why Diversity Indices Exist

If we want a number that responds both to how many species are present and how evenly individuals are distributed among those species, we need more than raw richness.

That is the motivation behind diversity indices such as Shannon or Simpson measures.

For example, one form of the Shannon index is:

H = −Σ pᵢ ln(pᵢ)

where pᵢ is the proportion of observed individuals belonging to species i.

The formula is less important here than the reason it exists.

Two communities can contain the same number of species while distributing individuals very differently. A diversity index tries to preserve some of that distributional information.

But an index is still a compression.

Two ecologically different communities can sometimes produce similar index values.

So the index should never be mistaken for the ecosystem itself.

11. The Species–Area Problem: Bigger Samples Usually Find More Species

If we survey 10 square metres of forest, then 100 square metres, then one hectare, the larger search area generally creates more opportunities to encounter different species.

This means comparisons must control for sampling area or model the area effect explicitly.

A common ecological form is the species–area relationship:

S = cAᶻ

where S is species richness, A is area and c and z are fitted parameters under a particular model and context.

The equation should not be treated as a universal machine that predicts Bukit Timah biodiversity from area alone. Habitat quality, history, isolation, edge effects, taxonomic group and sampling design all matter.

The useful lesson is structural:

Observed richness depends partly on how much of the world you looked at.

12. Scale Changes the Question

Consider four possible questions:

  • What species occur in one 100 m² plot?
  • What species occur across a 2-hectare permanent plot?
  • What species occur across Bukit Timah Nature Reserve?
  • What species occur across Singapore?

The word “biodiversity” appears in every question.

The statistical object changes each time.

Ecologists often distinguish local diversity from diversity across larger landscapes because turnover between sites matters. If two plots contain different species, the landscape can have greater overall diversity even if each local plot has modest richness.

This is a general modelling lesson:

Before analysing a system, declare the scale at which the system exists.

13. Convenient Data Can Be Biased Data

Imagine surveying butterflies only along the busiest trail at noon.

The observations might be real.

The inference to “butterflies in the Reserve” might still be weak.

Why?

  • The trail edge may have different light and vegetation from forest interior.
  • Noon conditions may favour some species and reduce activity in others.
  • Human disturbance may change presence or detectability.
  • Some microhabitats may never be sampled.

This is sampling bias: the data systematically over-represent some parts of the target population or process.

More biased observations do not automatically remove the bias.

Ten thousand convenient observations can be less representative than one thousand well-designed observations.

This principle matters far beyond ecology. It matters in opinion polls, medical studies, machine-learning datasets, school surveys and financial backtests.

14. The Observer Is Part of the Measurement System

Two observers can walk the same forest and record different data.

One may recognise a bird call the other misses. One may identify a plant to species while another records only the genus. One may measure trunk diameter at a slightly different height. One may overlook a cryptic insect.

Scientific protocols exist partly to reduce these differences.

  • standard measurement rules;
  • training;
  • specimen or photographic verification;
  • repeat observations;
  • calibrated instruments;
  • quality-control checks;
  • consistent taxonomic references.

The mathematical principle is that measurement error has structure.

Some error is approximately random. Some is systematic. The distinction matters because averaging repeated measurements can reduce some random noise but cannot automatically remove a persistent systematic bias.

15. Random Error and Systematic Error Behave Differently

Suppose a trunk-diameter instrument has small random fluctuations around the true value.

Repeated measurements may scatter above and below the truth. Averaging can help stabilise the estimate.

Now suppose the instrument is miscalibrated and always reads 2 millimetres too high.

Taking 100 measurements does not solve the calibration error. It may simply produce a very precise estimate that is consistently wrong.

Precision is not the same as accuracy.

This distinction is central to serious data work.

16. Uncertainty Should Travel With the Estimate

Students are often trained to feel that an answer with a decimal point is more scientific than an answer with a range.

In statistical inference, the opposite can be true.

If we estimate a population quantity from a sample, a responsible result often includes uncertainty.

For a simple estimated mean, one familiar form is:

estimate ± margin of error

or a confidence interval constructed under stated assumptions.

Biodiversity estimation can require more specialised methods, but the philosophical discipline is the same:

Do not report more certainty than the design and data have earned.

A range is not intellectual weakness.

Sometimes it is the most honest form of precision available.

17. Long-Term Monitoring Turns a Photograph Into a Film

A one-time survey answers, roughly, “What did we observe now?”

A permanent plot repeatedly surveyed through time can answer a richer family of questions.

  • Which trees survived?
  • Which died?
  • Which recruited into the measured population?
  • Which species increased or declined?
  • How did size distributions change?
  • How did the forest respond to storms, drought, edge effects or other disturbances?

This is why Bukit Timah’s Forest Dynamics Plots matter. NParks notes that the two permanent 2-hectare plots are re-censused regularly so researchers can understand how the forest develops over time.

Mathematically, time introduces another dimension.

Instead of a dataset indexed only by individual and location, we now have repeated observations:

observation = f(individual, location, time)

The forest becomes dynamic rather than static.

18. A Baseline Is a Reference System for Change

To say that biodiversity “declined” or “increased”, we need a comparison.

A baseline provides that reference.

But baselines are not neutral merely because they are old.

If the first survey occurred after substantial historical habitat loss, then “no change from baseline” does not mean “restored to original condition”. It means no detected change relative to that selected starting point.

This mirrors the elevation lesson from The Mathematics of Bukit Timah Hill: a value means something only relative to its reference system.

Datum in surveying. Baseline in monitoring. Benchmark in education. Index year in economics.

Different fields, same mathematical habit.

19. Detecting Change Means Separating Signal From Natural Variation

Suppose a species is recorded 27 times in one survey and 23 times in the next.

Has the population declined?

Possibly.

But the count difference alone does not prove it.

  • Detection conditions may differ.
  • Sampling effort may differ.
  • The species may naturally fluctuate.
  • Observers may differ.
  • The sampled locations may differ.
  • Random variation may account for some of the difference.

A change-detection problem therefore needs an expected-noise model, repeated observations or another defensible comparison framework.

The larger lesson is important for students:

Difference is not automatically evidence of meaningful change.

20. Finding a New Species Does Not Mean the Species Just Arrived

NParks announced in 2020 the scientific description of Nervilia singaporensis, an orchid native and endemic to Singapore found in Bukit Timah Nature Reserve. The agency highlighted that the discovery showed there was still unknown biodiversity to find even in heavily studied, highly urbanised Singapore.

Mathematically, this distinction matters.

Discovery date ≠ arrival date.

A species can exist before it enters the dataset.

That creates what we might call an observation boundary. The biological system and the knowledge system are not identical.

Reality can contain information that our current dataset does not.

This is an extremely important idea in science, statistics and artificial intelligence.

21. Wallace’s 700 Beetle Species: A Lesson in Sampling Effort and Productivity

Alfred Russel Wallace famously wrote that in about two months in the Bukit Timah area he obtained no fewer than 700 beetle species, many new to him and to science at the time, from a remarkably productive patch of jungle.

The number is historically striking.

It should not be converted casually into a modern estimate of “the number of beetle species in Bukit Timah”.

Why not?

  • Wallace’s collection methods and effort had a particular historical context.
  • The landscape was different.
  • Taxonomic classification has changed.
  • His sampling area and target groups were not a modern random sample of the present reserve.
  • Collection counts reflect detectability and collector effort as well as underlying diversity.

The historical observation is valuable.

The inference has to respect the design that produced it.

For the historical owner, read Alfred Russel Wallace in Bukit Timah | Beetles → Observation → Natural Selection → Scientific Method.

22. “More Than Half of Singapore’s Native Plant Species” Is a Ratio With a Story Behind It

NParks has noted that Bukit Timah Nature Reserve occupies only about 0.23% of Singapore’s land area yet is home to more than 50% of the country’s native plant species.

That statement is powerful because two percentages are being compared.

  • share of national land area: about 0.23%;
  • share of recorded native plant species: more than 50%.

But these are not the same type of percentage.

One is a spatial fraction. The other is a species-coverage fraction.

Dividing 50 by 0.23 and declaring the Reserve “217 times more biodiverse per unit area” would be a seductive but potentially misleading simplification. Species richness does not scale linearly with area, species overlap matters, and the national denominator involves habitats very different from rainforest.

The headline comparison tells us the Reserve contains an extraordinarily concentrated share of native plant diversity.

It does not license every ratio we can construct from the two percentages.

Numbers can be individually correct and still form a bad ratio.

23. Missing Data Are Not Automatically Zero

Suppose a field sheet has no bird count for Plot 17 on one morning.

Does that mean zero birds were observed?

Not necessarily.

Maybe the survey was cancelled by rain. Maybe the recorder failed. Maybe the observer never visited. Maybe the data were lost.

Statistically:

missing ≠ zero

Treating missing values as zero can create an artificial decline.

The reason data are missing matters too. If observations are more likely to be missing during heavy rain, and rain also affects animal activity, the missingness is connected to the phenomenon being studied.

What looks like an empty spreadsheet cell may therefore be part of the model.

24. Correlation in a Forest Does Not Automatically Tell Us the Cause

Suppose plots with greater canopy openness also show more individuals of a particular plant species.

The association may be real.

It does not by itself prove that canopy openness caused the increase.

  • Both may be related to disturbance history.
  • Soil moisture may affect both.
  • Elevation may confound the pattern.
  • The plant itself may alter local structure.
  • The association may differ across forest types.

Statistics can reveal structure in data.

Causal claims require stronger reasoning about mechanism, design and alternative explanations.

25. Every Biodiversity Number Sits Inside a Model

Consider the statement:

“There are over 140 bird species in Bukit Timah Nature Reserve.”

Even a straightforward inventory number carries hidden definitions.

  • What time period counts?
  • Are migrants included?
  • Are accidental visitors included?
  • How are taxonomic revisions handled?
  • Must a species breed there, or merely be recorded there?
  • What evidence qualifies as a valid record?

The point is not to distrust official numbers.

The point is to understand why trustworthy numbers come with definitions and methods.

A number becomes scientifically useful when we know what process produced it.

26. The Same Rainforest Can Teach Mathematics From Primary School to JC

Primary school

Students can classify observations, count categories, compare frequencies, use bar graphs, reason about fractions and percentages, and learn the basic difference between “what we saw” and “everything that exists”.

Secondary 1–2

Students can compare samples, discuss bias, calculate proportions, examine sampling methods, use scatter plots and distinguish association from certainty.

Secondary 3–4

Students can work with probability, expected values, distributions, cumulative-frequency ideas, correlation and the assumptions behind repeated observations.

Additional Mathematics

The focus can shift toward functions and rates: species accumulation curves, fitted relationships, change through time and how local behaviour differs from global summaries.

JC Mathematics

Students can discuss probability models, sampling distributions, confidence, hypothesis testing, regression, model assumptions and why observational data do not automatically establish causation.

For the school-stage owner map, use the Singapore Mathematics Curriculum Overview | Primary 1 to JC.

27. The Teaching Lesson: Do Not Give Students Clean Data Too Early

Textbook statistics often begins after the difficult decisions have been removed.

The table is complete.

The categories are already defined.

No observations are missing.

No measurement is uncertain.

No sampling frame is biased.

No observer disagrees.

That is useful for learning a technique.

It is not enough for learning how data work.

A stronger Mathematics education occasionally gives students the mess first:

  • Where should we sample?
  • What counts as one observation?
  • What if we miss a species?
  • What if two observers disagree?
  • What if the sample is concentrated near a trail?
  • What does a zero mean?
  • How should we compare two samples of different sizes?
  • How certain are we?

Then the formula has a reason to exist.

28. The Rainforest Lesson Also Applies to AI

A model trained on data learns from the observations available to it.

If the dataset oversamples one group, misses rare cases, contains systematic labelling error or reflects only one environment, the model can inherit those limitations.

Bukit Timah biodiversity sampling gives a concrete analogy.

  • A convenient trail sample is like a convenient training dataset.
  • A rare undetected species is like an underrepresented edge case.
  • Taxonomic misidentification is like label noise.
  • Repeated censuses are like monitoring model drift through time.
  • Changing habitat is like a changing data-generating process.

The analogy has limits, but the mathematical discipline carries across:

Never confuse the dataset with the world that produced it.

29. What Can the Data Actually Tell Us?

A responsible answer has several layers.

The data can describe observations

Researchers can report what was observed, where, when and under which method.

The design can support inference

If the sample is appropriate, observations can support estimates or comparisons beyond the exact sampled points.

Repeated data can support change detection

Permanent plots and consistent protocols can reveal trends through time.

Models can quantify uncertainty

Probability and statistics allow us to represent what is not perfectly known rather than hiding uncertainty behind a single number.

But the data cannot escape the sampling process

Every inference remains connected to the way observations were generated.

30. What This Article Does Not Claim

  • It does not claim that the two 2-hectare Forest Dynamics Plots alone represent every habitat in Bukit Timah Nature Reserve.
  • It does not treat NParks’ species inventory figures as perfect counts of every individual organism.
  • It does not infer population size from species richness.
  • It does not treat non-detection as proof of absence.
  • It does not assume repeated observations are independent without justification.
  • It does not treat a biodiversity index as a complete description of an ecosystem.
  • It does not turn historical Wallace collections directly into modern abundance estimates.
  • It does not treat area fraction and biodiversity fraction as linearly interchangeable.

These are not caveats added to weaken the argument.

They are the argument.

Good statistics is not the art of making uncertainty disappear. It is the discipline of making uncertainty visible enough to reason with.

Evidence and Official Sources

The Bukit Timah facts in this article were checked against current NParks and Singapore Botanic Gardens research material. Numerical probability examples are illustrative unless explicitly identified as observed field data.


So, How Many Species Are Really in Bukit Timah?

The most mathematically mature answer is not one heroic integer.

We have observed inventories. We have permanent plots. We have transects. We have historical collections. We have repeated censuses. We have new discoveries. We have taxonomic revisions. We have detection limits. We have uncertainty.

That does not mean we know nothing.

It means knowledge has structure.

The better the sampling, the clearer the definitions, the stronger the repeatability and the more honest the uncertainty, the stronger the inference becomes.

Bukit Timah rainforest teaches a statistical rule worth carrying everywhere:
what you observed is data; what you say about the whole world from those observations is inference.

Return to: What about Bukit Timah? — Live Hub · How Bukit Timah Nature Reserve Works · Alfred Russel Wallace in Bukit Timah · The Mathematics of Bukit Timah Hill · Singapore Mathematics Hub

Discover more from Bukit Timah Tutor

Subscribe now to keep reading and get access to the full archive.

Continue reading