Method & source
Where the data comes from
All figures on this site come from the public API behind rausgegangen.de. This site is not official, and it is not affiliated with or endorsed by Rausgegangen.
The site only publishes aggregate statistics like counts, shares, medians and distributions. It doesn't republish their catalogue: there is no bulk export, and there are no event descriptions or images. When a specific event is named, its title links to the event's page on rausgegangen.de.
The collection window
The API only returns events from today onwards. Past dates return nothing, so older events can't be collected after the fact. The archive only contains what was collected while it was current.
The crawler goes through each city one day at a time. Only the days it went through are complete, so every count on this site is limited to that window. The database also contains events further in the future, because fetching one event returns all of its future dates. Those are left out of the charts: they only cover the events the crawler happened to fetch, so including them would make the number of events look like it drops off over time.
Corrections
Times are converted to local time. Start times are stored as timestamptz
and the database runs in UTC. Reading the hour directly would give UTC and shift
every time by one or two hours depending on daylight saving time, so all times
are converted to Europe/Berlin first.
All-day listings are left out of time charts. Listings in the "Jederzeit verfügbar" category, like walking tours, escape rooms and city rallies, are stored as running from 00:00 to 23:59 instead of having a real start time. Counted by hour, they would all land on the same hour and create a fake spike. They are left out of anything based on the time of day, but still counted in totals and categories.
Prices
Prices come as text, not numbers: "19,00 bis 44,00 €", "Eintritt frei",
"Ticketshop", "Preis an AK". Each price is first classified as free, fixed,
range, donation or unknown, and only the ones with numbers are used in the price
charts. About one in twenty listings has no number at all. Counting those as
free would pull every average down, so they are left out.
The map
Coordinates come from the listings themselves. They are only included when a listing comes from a search result, so listings found through their parent event have none and are missing from the map. The number of missing listings is shown below the map.
Two problems in the source data are fixed before plotting. A few venues are stored at latitude 0, longitude 0, which would put them in the Gulf of Guinea, and at least one venue has its latitude and longitude swapped in the API response. Both are fixed during aggregation, so every chart uses the corrected positions.
Districts are assigned in the browser by checking which district polygon each venue falls into. Venues outside all districts, like a festival site outside the city, are shown on the venue map but not counted in the district totals.
The district outlines come from OpenStreetMap and are only there for orientation. They are simplified and stored in the repository, so building the site doesn't depend on an external service.
Venues, organizers and artists
In the API, venues, organizers and artists are all the same type of object with
a type field, so a listing's venue and its organizer are looked up through the
same endpoint. Venue figures count dated listings at a venue, not distinct
events. A museum showing one exhibition every day for a month counts as a month
of listings.
How the figures are built
The crawler, database schema, aggregation and this site live in one repository. The figures are generated by a data loader that queries the database directly whenever the site is built.