An image dataset of horse memes. Screenshotted by hand and collected by machine.
Why was the dataset created, and by whom?
The first image was taken on 26 June 2025. The last screenshot on 4 April 2026. Between them lie 282 days. On 28 of them I took screenshots. On 19 August 2025 there were 132.
The images show horses. Otherwise they have nothing in common.
I did not choose them alone. Instagram decides what appears in the feed; I decided what stays. The dataset is the record of both decisions, and it cannot be split into two parts.
Compiled by Stephan Nachreiner, class Networked Materiality, Academy of Fine Arts Nuremberg. No commission, no funding.
What is a data point, how many are there, and what is missing?
A data point is a screenshot with four attributes: date, original format, pixel dimensions, file size. Two were added later: the image text, read by OCR, and a motif class that I assigned.
For 125 images the date comes from the EXIF data of the capture. For 243 it comes from the file date. The difference is not decorative. A file date changes when a file is copied; an EXIF date does not. For 243 images I do not know for certain when I took them.
People appear in this dataset twice: in the images and behind them. Who made the images is known for 26 entries. For 342 it is not. The account name lay outside the crop I chose. I wanted the image. I cut away the name without thinking about it.
Among the 368 screenshots the largest class is called no text. It holds 142 images. That is 39 percent. A class that absorbs more than a third of a collection orders nothing. It only records what I did not want to decide while sorting. The images collected by machine have no class yet.
How was it collected, over what period, and with whose consent?
With an iPhone. Two buttons at once. No script, no scraper, no interface.
Collected while looking, not by plan. Hence the distribution: 252 images on two evenings, 116 on the remaining 26 days together. Anyone who draws conclusions about horse memes in general from this is drawing conclusions about two evenings in the summer of 2025.
There is no consent from the authors. I did not ask for it. That is the norm for image datasets taken from the web, and it is the point at which Adam Harvey's Exposing.ai begins. The difference is the number. With 1891 images I can name each one. With five billion no one can.
From 27 September 2026 a script was added. It queries twenty sources: meme sites, fediverse timelines, Reddit, image archives. six of them delivered. 1523 images came in this way. The script is part of the dataset.
The machine collection differs from the first in one respect. Where a platform names an account, it is recorded. For 723 of the 1523 images, authorship is documented as a result. With the screenshots it was lost. Here it survives if the platform names it. Mastodon and Reddit do. Imgflip, Know Your Meme and Cheezburger do not. Same collector, same interest. Only a different tool, and the tool helps decide what remains.
Of the 1523 images collected by machine, 675 are memes. 848 are not. The machine does not know the difference.
The reason lies in the search terms. Under the tag horse, a meme site shows a joke and a photo platform shows a horse. The query is the same. Of 600 images from Mastodon timelines, twelve were memes: two percent. The rest was equestrian photography, bronze sculpture, racecourse advertising, book covers and a great deal of My Little Pony.
I looked at all 1523 one by one and decided. Four verdicts: meme, not a meme, cartoon pony, unclear. There is no rule by which someone else would reach the same result. Whether a horse on a car roof is a meme or an accident photo is not a property of the image. It is my decision, and it is recorded with every image.
Rejected images were not deleted. They remain in the dataset, scaled down, with the reason for their exclusion, and can be shown in the field with show rejects. The silhouette shows 1043 images. The dataset contains 1891. The difference is not an error. It is the part that was decided on.
What was changed in the raw material, and what was lost?
254 images were cropped. Removed: status bar, clock, account name, the Follow button, the line Liked by … and others, the comment bar, the mute button, the dots marking an image's position within a post.
This crop is the most consequential intervention of the project. It turns a post into an image. An account into a blank. A second, automated pass found remnants of the interface in 42 images and removed them; the index card states what was cut in each case. Around 40 captures do not come from Instagram but from Reddit, Pinterest or search results. Their interfaces are partly still in the image.
For 26 images the account name could be recovered from the uncropped originals. For 342 it could not. The originals are fully preserved.
The image text comes from text recognition and is uncorrected. For 10 images it finds nothing, because there is none. I assigned the motif classes in a single pass, alone, without a second opinion, and invented the classes while sorting.
The images are laid out in the outline of George Stubbs' Whistlejacket (1762). Where the painting is light, lighter memes lie; where it is dark, darker ones. From a distance the painting's colour settles over them. Up close it falls away. The head has been redrawn in profile. In the painting it turns towards the viewer.
What has the dataset been used for, and what should it not be used for?
So far, as material for the following works.
Textile work. The material acquires weight and measure.
link to followPhotographic work around the blanket.
link to followThe place where the collecting happened.
link to followWhat it is not suited for: training data with any claim to being representative. It reflects one feed and a handful of platforms, in the German-speaking world, in 2025 and 2026. The classes come from one person. For the screenshots, authorship is 93 percent erased; across the whole collection, 60 percent. Whoever trains a model on it trains on my taste.
Under what licence is it distributed, and who maintains it?
I cannot grant a licence. The rights to the images lie with others, for 342 of them with people unknown. What is passed on is the work of collecting: the selection, the order, the description. That is less than a licence. It is more than most training datasets disclose.
The collection grows in two ways: screenshotted by hand and found by machine. Every image records which it was.
Anyone who recognises their own image and wants it removed will have it removed. No reason, no questions. It is the only form of consent that can still be established afterwards. Write to address to follow.
Timnit Gebru et al., Datasheets for Datasets (2018/2021) — the form of this document.
Mimi Onuoha, On Missing Data Sets (2016) — on the 342 blanks.
Hito Steyerl, A Sea of Data (2016) — on what falls out during cleaning.
Everest Pipkin, On Lacework (2020) — on looking at every single entry.
Adam Harvey, Exposing.ai — on consent and provenance.
Martin Herbert, Tell Them I Said No (2016) — on refusal, here carried out without asking.