Size First, Then Detail

Archaeobotany has a staffing problem. Identifying a carbonized seed under a microscope takes years of specialist training, and there are never enough trained people for the volume of material sitting in storage. A team at Shandong University and Lingnan University has built a system that does the first pass automatically, and the way they got it to work is more interesting than the fact that it does.

Start with what the specialist is actually doing. The published identification criteria for these seeds are, to a striking degree, ratios. Barley's embryo occupies about one-third of the length of the grain. Rice's runs to about one-fourth, set to one side. Wheat's is nearly circular and again about a third. Where a species can't be separated by outline — and charring flattens outline badly — the call comes down to what fraction of the seed one structure takes up, and how big the whole thing is in millimetres.

A photograph does not carry that information. An image classifier trained the ordinary way sees shape, texture and colour, scaled to a fixed input size, with absolute dimension discarded before training begins. This is not an oversight. It is what makes the technique general: a cat is a cat at any distance. For charred seeds it removes the single most reliable cue in the field.

The team put it back, twice. Every specimen was photographed at a fixed magnification, so true relative size survives into the image. Then the network was built to read that scale signal first, before judging any fine detail. The paper states the reasoning plainly: archaeologists sort approximately by size, then decide the species on features like the embryo region. The architecture is a transcription of that two-step procedure.

INTERACTIVE · drag the slider

What the photograph carries

embryo ratio: retained absolute size: retained

Seed dimensions are typical values for charred archaeological grain, drawn schematically. Embryo fractions follow the published identification criteria cited in Xing et al., npj Heritage Science (2026).

It worked. Against twenty-eight established classification methods, the best of which reached 84.9%, the new system reached 90.2%. The authors are direct about why the others fell short — none of them models size explicitly.

So the gain did not come from a more capable general-purpose model. It came from building a narrower one, from handing back a piece of domain knowledge the general approach had no way to represent.

That is the finding. Here is the part the coverage left out.

Accuracy is the wrong number to lead with, and the paper's own tables show it. The dataset is severely unbalanced: foxtail millet supplies more than 1,700 images, while peach stones supply forty-five, three of which sit in the test set. On a distribution like that, a model can post a high score by being reliable about the common categories and hopeless about the rare ones — which is what happens. The F1 score, which does not permit that, is 77.5%. On peach stones, precision is 33%: one call in three. The authors say in their conclusion that the problem is unsolved and flag it as future work.

The headlines called this the world's first AI archaeological system. The paper's title calls it a baseline model, and a baseline is a thing you publish in order to be beaten. The team has since opened it as a public challenge. What archaeobotany acquired this summer is not a machine that identifies seeds. It is a scoreboard, and a definition of the problem worth scoring.

That definition is where the open question sits. The dataset began with twenty-three categories and shipped with seventeen, the rest cut as too damaged to train on. Defensible one at a time — but together they fix what counts as an identifiable seed for every model that competes here from now on, and the categories that got cut are the ones that most needed the help. A benchmark does not only measure a field's progress. It selects the direction of it. The technical literature files that risk under long-tailed recognition, and it is a live problem there, not a settled one.

Sources: Xing, R., Cong, R., Wu, Y., Wang, C., Tang, Z., Wang, F., Wu, H. & Kwong, S. "Towards ancient plant seed classification: a benchmark dataset and baseline model." npj Heritage Science (2026). DOI 10.1038/s40494-026-02736-9. Preprint: arXiv:2512.18247, Dec 20 2025. Lingnan University press release, Jul 30 2026.
Previous
Previous

The Poison Factory

Next
Next

The Physicist Who Asks What Can’t Happen