Method & source

Where the data comes from

All figures on this site come from the public API behind rausgegangen.de. This site is not official, and it is not affiliated with or endorsed by Rausgegangen.

The site only publishes aggregate statistics like counts, shares, medians and distributions. It doesn't republish their catalogue: there is no bulk export, and there are no event descriptions or images. When a specific event is named, its title links to the event's page on rausgegangen.de.

The collection window

The API only returns events from today onwards. Past dates return nothing, so older events can't be collected after the fact. The archive only contains what was collected while it was current.

The crawler goes through each city one day at a time. Only the days it went through are complete, so every count on this site is limited to that window. The database also contains events further in the future, because fetching one event returns all of its future dates. Those are left out of the charts: they only cover the events the crawler happened to fetch, so including them would make the number of events look like it drops off over time.

Corrections

Times are converted to local time. Start times are stored as timestamptz and the database runs in UTC. Reading the hour directly would give UTC and shift every time by one or two hours depending on daylight saving time, so all times are converted to Europe/Berlin first.

All-day listings are left out of time charts. Listings in the "Jederzeit verfügbar" category, like walking tours, escape rooms and city rallies, are stored as running from 00:00 to 23:59 instead of having a real start time. Counted by hour, they would all land on the same hour and create a fake spike. They are left out of anything based on the time of day, but still counted in totals and categories.

Prices

Prices come as text, not numbers: "19,00 bis 44,00 €", "Eintritt frei", "Ticketshop", "Preis an AK". Each price is first classified as free, fixed, range, donation or unknown, and only the ones with numbers are used in the price charts. About one in twenty listings has no number at all. Counting those as free would pull every average down, so they are left out.

The map

Coordinates come from the listings themselves. They are only included when a listing comes from a search result, so listings found through their parent event have none and are missing from the map. The number of missing listings is shown below the map.

Two problems in the source data are fixed before plotting. A few venues are stored at latitude 0, longitude 0, which would put them in the Gulf of Guinea, and at least one venue has its latitude and longitude swapped in the API response. Both are fixed during aggregation, so every chart uses the corrected positions.

Districts are assigned in the browser by checking which district polygon each venue falls into. Venues outside all districts, like a festival site outside the city, are shown on the venue map but not counted in the district totals.

The district outlines come from OpenStreetMap and are only there for orientation. They are simplified and stored in the repository, so building the site doesn't depend on an external service.

Venues, organizers and artists

In the API, venues, organizers and artists are all the same type of object with a type field, so a listing's venue and its organizer are looked up through the same endpoint. Venue figures count dated listings at a venue, not distinct events. A museum showing one exhibition every day for a month counts as a month of listings.

How the figures are built

The crawler, database schema, aggregation and this site live in one repository. The figures are generated by a data loader that queries the database directly whenever the site is built.