Cast demographics is the aggregate picture: averaged across 4,361 titles that have a researched cast, what is the typical gender and racial composition of the people on screen, and how do leads differ from the broader ensemble? This is the catalog’s widest-angle lens. Where the swap and representation measures zoom in on individual characters, this one steps back and asks what the average looks like.
The headline gender figures are the most reliable, because gender is the most consistently researched attribute. On average, leading roles across the catalog skew male: roughly 52 percent of lead billing is male against 31 percent female, with the remainder distributed across non-binary, unknown, or unresolved entries. That gap is one of the clearest aggregate signals in the dataset.
How WhatSphere measures this
For each title, the cast is tallied into percentage shares by gender and by racial or ethnic group, computed separately for leads and for the full cast, and the director’s gender is recorded. Those per-title percentages are then averaged across the catalog. The averages describe central tendency, not any individual title; a single film can sit far from the mean in either direction. Every figure here is generated by WhatSphere’s scoring model, not by a critic’s judgement. Scores describe who a title appears to be made for and how its content is composed; they are not ratings of quality, and a higher or lower number is never a verdict on whether a title is good.
What the data shows
The lead-gender skew toward male billing is the firmest finding. Behind the camera, the same picture is more pronounced: of the titles with a recorded director gender, 1,019 have a male director and 52 a female one. Reporting on-screen and behind-camera figures together keeps a film with a balanced cast and a single-gender directing slate from being summarised by either number alone.
Racial composition is reported group by group rather than as a single index, and it carries a heavy caveat described below. Among the groups the model identifies, White is the largest recorded share of average cast, with the smaller groups trailing well behind. These shares are percentages of the cast that has been positively identified — not of the entire cast — which is why they should be read as a floor on each group rather than a full accounting.
Caveats and limitations
The racial-composition figures carry the single largest limitation in the dataset: for most cast members, race or ethnicity has not yet been positively researched and is recorded as unknown. That means the per-group percentages understate every identified group, because a large share of each cast is sitting in the unknown bucket rather than being attributed. The gender figures are far more complete and should be weighted accordingly. Averages also hide spread — a mean near the middle can come from balanced titles or from a mix of opposite extremes, so the distribution matters as much as the central value.
Why it matters
Claims about who is and is not on screen are made constantly and quantified rarely. A consistent, catalog-wide method for tallying lead and ensemble composition — with the on-screen and behind-camera numbers side by side and the data-completeness limits stated openly — gives those claims something to be checked against. Read together, the gender split, the lead-versus-ensemble cut, and the behind-camera breakdown give a fuller picture than any single percentage can, and the open treatment of the unknown-race bucket means the figures can be trusted for exactly what they claim rather than oversold for what they cannot yet measure. The titles linked below are the ones with the most divergent cast compositions, useful as concrete illustrations of how far individual films can sit from the catalog average.