How did the psychology of gender go from measuring brain sizes to developing personality inventories that could classify someone as “androgynous”? The journey spans less than a century, but the conceptual leap is enormous. From the late 1800s through the early 1980s, researchers moved from viewing gender as a fixed biological fact to recognizing it as a complex, multidimensional aspect of personality. Understanding this history is not just an academic exercise – it reveals how scientific assumptions shape what questions get asked, what gets measured, and who gets pathologized in the process.
Table of Contents
- Early studies on sex differences in intelligence (1894-1936)
- Havelock Ellis and the biological framing
- Lewis Terman and intelligence testing
- Masculinity and femininity as personality traits (1936-1954)
- The Terman-Miles Attitude Interest Analysis Survey
- The MMPI masculinity-femininity scale
- What this era got wrong – and why it mattered
- Androgyny and sex typing (1954-1982)
- The conceptual groundwork
- Sandra Bem and the BSRI
- The Personal Attributes Questionnaire (PAQ)
- Critiques and the limits of androgyny
- What this history reveals
Early studies on sex differences in intelligence (1894-1936)
The earliest chapter in the psychology of gender was written by researchers who were convinced that biological sex determined intellectual capacity. Early brain studies comparing mass and volume between the sexes suggested women were intellectually inferior because they have smaller and lighter brains – a claim that reflected the era’s broader social climate far more than any rigorous science.
The late 19th and early 20th centuries saw psychologists treating this question as scientifically urgent, partly because debates about women’s suffrage made the issue politically charged. If women were intellectually equal to men, the argument went, then denying them the vote was indefensible. Much of the early research in this period was therefore shaped – consciously or not – by those social stakes.
Havelock Ellis and the biological framing
British physician and psychologist Havelock Ellis was among the influential early voices who argued that anatomical differences between men and women extended to their intellectual and psychological capacities. His work helped establish the idea that sex differences in behavior and ability were rooted in biology, a framing that dominated the field for decades. Although Ellis’s specific claims have long since been discredited, his emphasis on biology as the primary explanatory framework proved remarkably persistent.
Lewis Terman and intelligence testing
Lewis Terman is best known for his work on IQ testing, including the Stanford-Binet Intelligence Scale. His 1916 revision of that scale contributed to early claims about sex differences in intelligence – yet, as Terman himself concluded in that 1916 study, “the intelligence of girls, at least up to 14 years, does not differ materially from that of boys.” That finding received far less attention than the popular narrative of male intellectual superiority, illustrating how science is never entirely separate from the cultural assumptions of its time.
Figures like Leta Hollingworth pushed back directly against biological determinism, arguing that women were not permitted to realize their full potential because they were confined to roles of child-rearing and housekeeping – an early and important recognition that social structures, not biology, might explain observed differences. By the early 20th century, the scientific consensus was shifting toward the view that sex plays no role in intelligence, even if popular belief lagged far behind.
The 1894-1936 period ultimately raised questions it could not adequately answer, largely because it lacked the conceptual tools to distinguish between sex as biology and gender as a social and psychological phenomenon. That distinction would take several more decades to develop.
Masculinity and femininity as personality traits (1936-1954)
By the mid-1930s, researchers began to shift their focus. Rather than asking whether men and women differed in raw intelligence, psychologists turned to personality – specifically, to the idea that masculinity and femininity were measurable psychological traits. This shift was significant: it moved gender out of the skull and into the self, treating it as something that could be assessed through attitudes, interests, and emotional responses.
The Terman-Miles Attitude Interest Analysis Survey
The pivotal instrument of this era was the Attitude-Interest Analysis Test, developed by Lewis Terman and Catherine Cox Miles and published in 1936. Often called the Masculinity-Femininity (M-F) Test, it was the first comprehensive attempt to quantify M-F traits. The test covered seven content areas: word association, inkblot association, information tests, emotional and ethical attitudes, interests, opinions, and introvertive response patterns.
The underlying logic was straightforward, if deeply flawed: find items that reliably differentiated men from women in a given population, then use those items to place any individual on a single masculinity-to-femininity continuum. Terman and Miles concluded that males showed a distinctive interest in exploit, adventure, outdoor and physically strenuous occupations, machinery, and science, while females demonstrated a distinctive interest in domestic affairs and aesthetic objects – findings that tell us more about 1930s American culture than about any timeless psychological reality.
This was the central problem with the entire approach: critics argued that scores heavily reflected the particular culture and time period in which they were constructed, emphasizing social differences of the most extreme sort. When the Terman-Miles test was applied in other cultures, only a fraction of its items differentiated the sexes – strong evidence that the test was measuring cultural conformity, not deep psychological difference.
The MMPI masculinity-femininity scale
The other major instrument of this period was the Mf (Masculinity-Femininity) scale embedded in the Minnesota Multiphasic Personality Inventory (MMPI), which became one of the most widely used personality assessments in clinical psychology. The MMPI’s Mf scale was developed partly by drawing on the Terman-Miles test, and it was originally used to identify men whose interests deviated from the male norm – a framing that pathologized gender nonconformity by linking it directly to psychiatric assessment.
Both instruments shared a critical theoretical assumption: that masculinity and femininity were bipolar opposites on a single continuum. A person could be more masculine or more feminine, but not both. This either/or framework would eventually be dismantled, but for nearly two decades it defined how psychologists understood and measured gender personality.
What this era got wrong – and why it mattered
Early M-F tests conflated several distinct variables: biological sex, sexual orientation, gender role, and personality. The assumption that these all moved together on a single scale had real consequences – it meant that a man with stereotypically feminine interests could be classified as psychologically deviant, not simply as someone with a broad range of traits. This period established measurement practices that would take decades to reform, but it also generated the data and critiques that made reform possible.
Androgyny and sex typing (1954-1982)
The third and most conceptually revolutionary period in the history of gender psychology began with a growing dissatisfaction with the bipolar model. If masculinity and femininity were truly opposite ends of one spectrum, then any increase in feminine traits logically required a decrease in masculine ones. But researchers increasingly suspected this was not how personality actually worked. The result was a paradigm shift: masculinity and femininity came to be understood as independent dimensions, not opposing poles.
The conceptual groundwork
Through the 1950s and 1960s, several researchers began questioning whether the bipolar assumption held up empirically. Studies found that many individuals scored moderately on both masculine and feminine measures, and that high scores on one did not reliably predict low scores on the other. This opened the theoretical door for the concept of androgyny – the possibility that a single person could authentically possess high levels of both masculine and feminine characteristics.
The late 1960s through the 1970s marked an important turning point in the field of gender research, coinciding with second-wave feminism and broader cultural reconsideration of rigid sex roles. Researchers began asking not just whether men and women differed, but whether those differences were healthy, adaptive, or socially imposed.
Sandra Bem and the BSRI
The most influential figure in this period was psychologist Sandra Bem. In 1974, she published the Bem Sex-Role Inventory (BSRI), a landmark tool that formally operationalized the two-dimensional model of gender. Rather than placing respondents on a single M-F spectrum, the BSRI measured masculinity and femininity as two fully independent scales.
The BSRI consisted of 60 personality traits: 20 stereotypically masculine (such as assertive, independent, dominant), 20 stereotypically feminine (such as warm, nurturing, compassionate), and 20 neutral filler items. Respondents indicated how well each item described themselves on a 7-point scale, with the masculinity score and femininity score calculated independently – meaning a person’s standing on one scale had no bearing on their standing on the other.
This design allowed for four distinct classifications: masculine (high M, low F), feminine (low M, high F), androgynous (high on both), and undifferentiated (low on both). Bem found that some individuals had balanced levels of traits from both scales, and she described those individuals as androgynous – a category that simply had not existed in the earlier bipolar framework.
Bem’s theoretical argument was that androgynous individuals enjoy a psychological advantage: by being able to draw on both instrumental (typically “masculine”) and expressive (typically “feminine”) capacities, they can respond more flexibly to a wider range of social situations. Androgynous persons were theorized to avoid the limitations of polarized identities, which might impair versatility in diverse social environments.
The Personal Attributes Questionnaire (PAQ)
Around the same time, Janet Spence and Robert Helmreich developed the Personal Attributes Questionnaire (PAQ), another instrument built on the two-dimensional model. Like the BSRI, the PAQ treated instrumentality (agency, self-assertion) and expressiveness (warmth, interpersonal sensitivity) as independent dimensions. Spence argued that the terms “masculinity” and “femininity” were actually inappropriate labels for the scales, since what they really measured were instrumentality and expressiveness – dimensions present in all people, not exclusive to one sex.
This critique pointed toward a deeper conceptual issue: even the improved two-dimensional model risked perpetuating the assumption that certain traits were inherently gendered. The PAQ’s framing – focusing on instrumentality and expressiveness as personality variables – was a step toward eventually decoupling these traits from gender labels altogether.
Critiques and the limits of androgyny
By the early 1980s, the androgyny model itself was facing serious criticism. Researchers pointed out that the BSRI’s items were chosen based on what 1970s American college students considered culturally desirable for each sex – a sample and a moment in time that could not represent universal psychological categories. Even the BSRI and PAQ models faced criticism for assuming that high scores on both M and F were inherently superior to low scores on both, when subsequent research suggested the optimal profile depended heavily on specific context and cultural expectations.
Additionally, increasing criticisms pertaining to the methodology and conclusions drawn from gender differences research resulted in a decrease in studies focusing on such differences by the 2000s. The field was moving toward more nuanced frameworks – ones that would eventually incorporate social construction theory, intersectionality, and a broader questioning of whether binary gender categories were adequate to the complexity of human experience.
What this history reveals
Tracing these three periods – from brain-size comparisons to bipolar M-F scales to the two-dimensional androgyny model – reveals something important: scientific frameworks are not neutral. Each era’s research tools encoded assumptions about what gender was, what it meant to deviate from norms, and whose experience counted as normal. The shift from viewing gender as a biological fixed point to treating it as a multidimensional psychological construct was not just a technical improvement in measurement. It reflected broader social transformations, including the women’s rights movement, evolving clinical ethics, and growing recognition that pathologizing gender nonconformity caused real harm.
The inventories developed in this period – particularly the BSRI and Bem’s subsequent gender schema theory – remain foundational references in gender psychology, even as the field has moved well beyond them. They are best understood not as final answers, but as crucial stepping stones toward a more honest and complex understanding of gender.
What do you think? If masculinity and femininity are really just culturally defined labels for instrumentality and expressiveness, should psychology stop using gendered terms for these traits entirely? And looking back at these three historical periods, which shift – from biology to personality traits, or from bipolar to two-dimensional models – do you think represented the more significant conceptual breakthrough?
References
- https://en.wikipedia.org/wiki/Sex_differences_in_intelligence
- https://collection.sciencemuseumgroup.org.uk/objects/co8237494/attitude-interest-analysis-test-also-known-as-masculinity-femininity-scale-material
- https://scales.arabpsychology.com/trm/masculinity-femininity-test/
- https://scales.arabpsychology.com/trm/masculinity-femininity-tests/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3131694/
- https://www.britannica.com/science/Bem-Sex-Role-Inventory
- https://grokipedia.com/page/Bem_Sex-Role_Inventory
- https://en.wikipedia.org/wiki/Sandra_Bem
Leave a Reply