Play Styles: Grounding, Design, and Validity
National Institute for Play
Version 1.1, August 2026
Public edition. This is the published edition of the rationale report. It contains the full theory, model, methodology, evidence, limitations, and privacy sections. Verbatim item wording, exact scoring formulas and constants, and the raw-data variable dictionary are summarized rather than reproduced, to protect the instrument. Researchers, reviewers, and prospective partners can request the complete codebook under a written agreement: inquiry@nifplay.org.
Summary
Play Styles is a self-report instrument that identifies how a person most naturally enters the state of play. A Play Style score chart visualizes how much of each Primary Style and corresponding Subtypes a person possesses based on their answers. The free assessment asks 34 items (18 Likert, 3 forced-choice pairs, 12 context sliders, and 1 six-item ranking) plus 1 unscored attention check and 4 optional demographic questions. The premium assessment includes all of those and adds 48 subtype Likert items, 6 within-style rankings, 6 dimension sliders, a barriers question, 4 unscored Play History items, and 1 optional private reflection — about 105 prompts in total. There are 6 primary Play Styles, 24 subtypes, and 7 dimensions of play.
Play Styles is grounded in four decades of play histories collected by the National Institute for Play, and in affective neuroscience, which places play among our seven primary emotional systems (seeking, caretaking, rage, lust, fear, panic, and play).
Every respondent answers the identical set of questions; the presentation order of the Likert items is randomized per respondent with a stable per-session seed. The instrument mixes four formats deliberately: Likert items for how much something describes you, forced choices and rankings for determining which option is most like you, and sliding scales for context. Each format fails differently, so using all four produces a sharper profile than relying on only one.
We have analyzed 20,966 completed assessments. Two key findings: first, the framework's opposite pairs were derived from the data rather than assigned by theory, and the empirical results aligned with the conceptual model — a derivation since re-run and confirmed on the full sample using raw composite scores. Second, people who play more report meaningfully lower burnout, and people whose childhood play was significant are more playful as adults.
We have not established test-retest reliability or factor structure, outlined in Part V. We have also chosen not to validate Play Styles against clinical or established psychological scales. It is a tool for self-insight and shared language, not a clinical or diagnostic instrument, and should be judged on internal coherence, reliability, and usefulness.
How to use this document
If you are curious about where Play Styles came from, read Parts I and II. This includes the theory, the model, and the six styles defined.
If you are deciding whether to bring this into an organization, read Parts I and II for theory, Part IV for what the data shows and its limits, Part V for what has not yet been established, and Part VI for how personal information is handled.
If you are an analyst or reviewer, Part III is the codebook: the structure of every block, how each contributes to a score, and how the dataset is organized. Verbatim item wording, exact scoring constants, and the variable dictionary are withheld from this public edition and shared under agreement.
Contents
Part I: Grounding. Purpose of the document, what Play Styles is and is not, the nature of play, theoretical foundations, and the limits of any personality framework.
Part II: The model. The six styles, 24 subtypes, 7 dimensions, structural opposites, and character strengths.
Part III: The instrument. The codebook: item banks, response scales, scoring logic, design rationale, validation plan, and dataset conventions.
Part IV: Evidence. The pre-launch simulation, the preliminary analysis of 3,469 responses, the first full psychometric evaluation of the item bank on 3,730 cleaned responses, how those findings shaped the report, and the limitations of the sample.
Part V: What is outstanding. What we have not yet established, in priority order, and what we have chosen not to pursue.
Part VI: Privacy and data protection. What we collect, why, how long we keep it, and how a person can retrieve or delete it.
PART I: GROUNDING
1. Why this document exists
This document draws the curtain back to share the theory behind Play Styles and the methodology, logic, and mechanism of the assessment.
There is valid criticism and skepticism across the personality assessment industry, and we understand and share the limitations of this assessment openly. Instruments have been sold with claims their evidence does not support, used for decisions they were never validated for, and defended with marketing rather than data.
Play Styles is a new instrument built on a longitudinal body of established play science work. This document states what we have and have not established, and the process we are following to ensure continued research, adaptation, and refinement of this assessment tool. Section 13 is a list of our own open questions, published deliberately.
2. Play Styles Overview
Play Styles is a self-report instrument that identifies a person's natural entry points into the state of play, expressed as a primary and secondary style, subtypes within those styles, and a profile across seven dimensions of play.
Play Styles is not:
- A clinical or diagnostic instrument. It does not assess mental health, and no result should be read as a diagnosis or a clinical recommendation.
- A hiring, selection, or promotion tool. It has not been validated for employment decisions and must not be used for them.
- A measure of ability, potential, or worth. There is no established best or worst Play Style or mix of styles.
- A validated psychological scale in the clinical sense. It is a self-insight tool, evaluated on internal coherence, reliability, and usefulness rather than on convergence with an established validated construct.
- A fixed type. Play Styles describes preferences, not categories. A person's results tend to hold steady over time, but a style does not determine behavior. Context, mood, environment, and choice all shape how a person acts on any given day, and two people who share a primary style can play very differently. A result describes where someone tends to find play, not what they will do.
Note: Most of the criticism of personality assessments comes from use outside the purpose they were built for.
Design rationale
Every respondent answers the same questions.
There is no adaptive branching, no mid-assessment decision tree, no item trimming, and no personalization of the question set based on how someone has answered so far. Two people taking Play Styles a year apart on opposite sides of the world see an identical instrument. The only variation is presentation order within the Likert blocks, which is shuffled per respondent with a stable per-session seed; item identity and scoring are unaffected.
This is deliberate. An adaptive assessment can feel more responsive and can sometimes reach a result in fewer questions. A static design keeps the research dataset rectangular, so every response can be compared with every other response across each item. It makes item-level psychometrics possible (reliability, discriminant validity, and factor analysis). And it prevents later items from being conditioned on earlier scores, which would quietly build the model's assumptions into the data used to test the model.
The practical effect is that every person who completes the assessment adds a complete row to the same table. A branching instrument with the same number of respondents would yield a fraction of the usable evidence.
We value the highest standards in data collection for research and chose a static assessment design to match that value over a potentially easier user experience for participants.
A combination of scales
We deliberately chose to use Likert Scale, Forced Choice, Ranking, and Sliding Scale questions. The variety in scales provides a range of data types and creates a variety of filters to assess preferences from different angles.
Each format asks a different kind of question. Likert items ask how much a statement describes you, which produces scores that can be compared across people. Forced Choice pairs ask which of two options is more you, a trade off that separates preferences a Likert scale would leave tied. Rankings extend that trade off across a full set. Sliding scales capture degree on context and state questions where a five point scale would be too coarse.
Each format also fails in a different way. Likert items are vulnerable to agreement bias and to flat profiles where every style scores alike. Forced Choice and Ranking produce results that describe a person against themselves rather than against other people. Sliding scales invite the midpoint. Combining all four keeps the score comparable across people while forcing the differentiation a single format would miss. Section 7.6 sets out the analyst facing detail.
3. Understanding play
The essence of play
Play is biologically hardwired into the subcortex, the most ancient part of the brain (Panksepp, 1998; Vanderschuren, Achterberg, & Trezza, 2016). The subcortical location means play is not just important for kids, it is important across our entire lifespan. Play deprivation can have severe consequences at all ages (Brown & Vaughan, 2009; Gray, 2011), yet our society often treats play as a waste of time instead of a core component of our public health. The National Institute for Play exists to challenge social misconceptions using scientific evidence, and Play Styles is the mechanism to help individuals and communities integrate play as a meaningful habit.
Defining play
A simple, universally accepted definition of play has long eluded researchers, scientists, and clinicians alike. Play is complex, multi-dimensional, and deeply personal. What one person experiences as play may feel entirely different to another. The range of experiences that fall under the play umbrella remains a topic of lively debate.
At the National Institute for Play, we define play on a spectrum between a state and a trait.
The characteristics of the play state include but may not be limited to:
- We lose track of time. It can feel like an alternate state of being: a temporary world where new rules apply and new possibilities emerge.
- We experience a diminished consciousness of self.
- We engage in it for its own sake, not for a particular outcome.
- It is flexible, adaptive, and can be spontaneous.
- It evokes powerful emotions such as joy, glee, and awe.
- It is self-reinforcing. It keeps us coming back for more and has long-term benefits.
The play trait characteristics include but may not be limited to:
- A cognitive disposition towards playfulness: a way of being and showing up in the world, choosing playfulness in what could otherwise be ordinary.
- Adaptability
- Creative problem solving
- Optimism
- Authentic set of values
- Humor
- Resilience
Deep and frequent engagement in the state can strengthen the trait over time, and playfulness in turn supports long-term cognitive, social, emotional, and physical health across the lifespan.
The complexity of play arises because it is fluid and deeply personal. What feels like play for one person may not for another, and what feels playful one day may not the next. Internal and external environments shape our access to the play state moment by moment. Two colleagues coding side by side may be having entirely different experiences: one immersed in playful flow, the other simply working. A small shift, a distraction breaking the flow or a favorite song sparking movement, can change who is in the play state and who is not. This is where Play Styles comes in. It identifies where a person's natural entry points to the state of play reside in everyday life, and gives them a toolbox for reaching those entry points more easily and more often.
Play cues. We can often recognize play through signals like laughter, smiles, gleeful expressions, banter, open body language, and playful gestures.
This distinction matters for the Play Styles instrument. Play Styles measures preferred entry points into the state of play. It is not a measure of the playfulness trait or of the subsequent character traits established through the play trait. Those are separate constructs, measured separately, and we report how they relate in our research analyses rather than treat them as the same thing.
Why play matters for human flourishing. Drawing on Self-Determination Theory (Ryan & Deci, 2000), humans have three basic psychological needs: autonomy, competence, and relatedness. Play satisfies all three. Through play we experience autonomy, the freedom to choose and express without external pressure. We build competence through challenges that stretch just beyond current ability. And we build relatedness through shared experience that deepens connection and fosters belonging.
Beyond those needs, play serves as a vehicle for self-expression, allowing us to externalize an inner world that words cannot fully capture. It supports meaning-making. It offers the nervous system respite through joy and laughter. And it cultivates the creativity and cognitive flexibility that underpin original thinking.
In a world that is experiencing loneliness, polarization, disconnection, and information overload, play offers a foundational, biologically hardwired pathway to flourishing.
4. Theoretical grounding
4.1 Play as a primary emotional system
Neuroscientist Jaak Panksepp's work in affective neuroscience (Panksepp, 1998; Panksepp & Biven, 2012) identified seven primary emotional systems in the mammalian subcortex: SEEKING, CARETAKING, PLAY, LUST, FEAR, PANIC, and RAGE. These systems are not learned. They are embedded in every human at birth as instinctive behaviors, and they evolve over time as we accumulate memories from lived experience.
Panksepp described three levels of processing:
- Primary processes. Subcortical emotional systems, the raw sparks. Produce an emotional response.
- Secondary processes. Learning and memory regions shape those sparks into habits. Produce learned responses and choices.
- Tertiary processes. Cortical thought, reflection, and symbolic meaning. Produce reasoned decisions.
A child meeting a dog for the first time illustrates all three. A sudden bark fires the FEAR system: heart racing, body frozen or fleeing. A wagging tail and an invitation to play fires PLAY or CARETAKING: giggles, curiosity, reaching out. Nobody teaches these responses. They are neural reflexes older than our species. Over time the child learns which dogs are safe, which is secondary processing. By around age six, the child begins making reflective decisions about each encounter, which is tertiary.
This is why we say we are built to play, and built by play.
4.2 From emotional systems to personality
Panksepp and Kenneth Davis argue that personality emerges in part from differences in the sensitivity of these neural systems (Davis & Panksepp, 2018): the strength or weakness of the signals emanating from them. Some people have more reactive FEAR circuits, others more reactive SEEKING or PLAY circuits. These built-in tendencies give each of us a distinct emotional signature, an affective personality.
These systems are action-oriented and continuously, subconsciously influence thought, perception, and action. Play is our joy system, and it motivates physical and social engagement.
An honest caveat on this grounding. Panksepp's affective neuroscience is well established as a theory of emotional systems. The step from "PLAY is a primary emotional system" to "individuals have identifiable, stable Play Styles" is a plausible inference, not a demonstrated neuroscientific finding. The neuroscience motivates the model and gives it a principled foundation. It does not by itself validate the six-style structure. What supports that structure is the body of play histories described next and the empirical work in section 9, and what will confirm or revise it is the research in section 13.
4.3 The play histories grounding
Play Styles is grounded in patterns derived from four decades and thousands of personal play histories collected by Dr. Stuart Brown and the National Institute for Play.
These personal play histories revealed stylistic play preferences: evidence of a biological design to play in unique personal ways. The variety of play choices people described reflects both an innate drive to play and the diversity of activities that can activate a play state. These preferences are most vivid in childhood and through the developmental years, and they persist into adulthood unless suppressed, overridden, or ignored.
When activated, Play Styles express intrinsic motivation. When nourished, they help shape personality, build a sense of authentic self, sustain engagement, drive purpose, and develop strengths of character. They are not trivial. They flow from evolutionarily embedded subcortical circuits that ground our affective life, and they encompass the interplay of emotion, cognition, and behavior in constant motion.
4.3.1 What is a play history?
A play history is a structured personal narrative interview developed by Dr. Stuart Brown over five decades of clinical and research practice. It begins in early childhood and traces the arc of a person's relationship with play across their lifetime including what drew them in, what brought them joy, what they gravitated toward when free to choose, and what happened to that relationship over time.
The play history is not a checklist. It is a conversation designed to surface patterns that the person themselves may never have consciously examined. Play histories have been taken in clinical settings, in research contexts, in workshops, in group settings, and in individual conversations with children, adolescents, and adults across a wide range of backgrounds, cultures, and life circumstances.
4.3.2 The question arc: categories and representative questions
Play histories are organized around several broad categories of inquiry. The questions below are representative of the kinds of questions asked within each category. They are not exhaustive, and the full question set has evolved over decades of practice.
1. Early childhood play memories
- What did you love to do as a young child when no one was telling you what to do?
- Look at photographs of yourself before age eight. What do they show you doing, and what does that evoke?
- What made you laugh out loud as a child?
2. Freedom and self-direction
- Think of times in your life when you have been completely free to do what you choose. What do you do?
- Are there activities or situations in which you feel entirely free to be yourself?
3. Flow and absorption
- Can you recall a time when you were so engaged in something that you looked up and far more time had passed than you expected?
- Have you ever forgotten other responsibilities because you were so absorbed in what you were doing?
4. Joy, lightness, and aliveness
- What activities or situations make you feel most alive?
- When do you feel light — not burdened — in what you are doing?
5. Objects, environments, and contexts
- Are there particular objects, materials, or physical environments you are consistently drawn to?
- Are there specific types of music, locations, books, or sensory experiences that reliably create a sense of satisfaction or pleasure?
6. People and social context
- Do you play differently alone than you do with others?
- Are there particular people or types of people in whose presence you feel most free to play?
7. The arc of play across a lifetime
- Were there periods in your life when play disappeared or was suppressed? What happened?
- What activities did you love as a child that you have stopped doing, and why?
8. Play in adult life
- What do you do now, if anything, that feels genuinely playful rather than obligatory?
- Is there something you have always wanted to try but have not yet given yourself permission to do?
4.3.3 How many play histories have been taken, over what timeframe?
Dr. Brown began taking formal play histories as a psychiatric resident in the 1960s, initially as a tool for understanding the developmental roots of his subjects' emotional lives. Over the following five decades, play histories were taken in clinical settings, research interviews, workshops, public presentations, and informal conversations. The total number of play histories is not officially documented, but is in the thousands.
Play histories have been taken with individuals ranging from children to adults in their nineties.
More recently, play histories have been taken in group settings and workshops associated with the National Institute for Play. Family members and close collaborators of the Institute have also contributed their own play histories, in some cases providing a multigenerational perspective on how play patterns develop and persist across generations.
BBC Soul of the Universe (1991). In 1991, Dr. Brown conducted play history interviews for the BBC documentary series Soul of the Universe, speaking with Nobel laureates and other distinguished figures about the role of play in their lives and work. These conversations offered a window into how play manifests at the highest levels of human achievement, and contributed early qualitative data to the growing pattern library that would eventually inform the Play Styles framework.
PBS The Promise of Play (2000). In 2000, Dr. Brown collaborated on the production of The Promise of Play, a three-part PBS documentary series. As part of that production, interviews were conducted with members of the public engaged in leisure and recreation, exploring what they sought, enjoyed, and preferred in their play. This informal body of material added a demographically diverse layer of play history data to the corpus.
Adolescent play histories: a high school research initiative (2020). In 2020, young research assistants affiliated with the National Institute for Play undertook an independent play history project with peers in a high school setting. Conducting structured play history interviews with fellow students, they gathered first-person accounts of play across childhood and adolescence including how it had been experienced, what forms it had taken, how it had evolved or disappeared under academic and social pressure, and what felt genuinely playful to the subjects.
The summary report of that research offers a rare adolescent-centered perspective on the play history methodology and its findings. Several themes emerged consistently across the interviews: the early dominance of physical and outdoor play in childhood, a marked contraction of play during the middle school years as academic and social pressures increased, the persistence of creative and imaginative play in private even as public play became less socially acceptable, and a genuine hunger among adolescents for more unstructured time and permission to play without purpose.
This adolescent dataset is a meaningful addition to the play history corpus for two reasons. First, it captures the play histories of a population (teenagers) whose relationship with play is potentially vulnerable to suppression and underrepresented in the broader literature. Second, it demonstrates that the play history methodology translates effectively across age groups and can be conducted by trained individuals beyond the clinical setting, supporting the scalability of the approach as a research tool.
The key findings from this summary report are consistent with the broader patterns that informed the Play Styles framework: play is deeply individual, it follows recognizable, categorizable patterns, it is shaped powerfully by environment and social permission, and it does not disappear with age.
4.3.5 How the patterns were established
The six Play Styles emerged from Dr. Brown's observation of recurring patterns across the play histories he collected. Certain themes appeared repeatedly: people who consistently gravitated toward storytelling and imagination; people whose play was fundamentally physical and kinetic; people who found their deepest engagement in making and building things; people whose play was social and relational at its core; people drawn to competition and the dynamics of winning and losing; and people whose play was rooted in ideas, humor, and the play of the mind.
These were not categories imposed on the data from the outside. They emerged from the histories themselves, from the accumulated weight of thousands of people describing, often for the first time, what play had meant to them across their lives. The six styles represent the most consistent and distinguishable patterns across that body of evidence.
The subtypes within each style reflect finer-grained variation observed within those broader patterns, including ways in which people who share a primary style orientation still differ meaningfully from one another in how that style expresses itself.
The adolescent play histories collected in 2020 provided an additional layer of confirmation: the same broad patterns observed across adult populations were recognizable in younger respondents, suggesting that play style preferences emerge early and persist across development even when their expression changes with age and circumstance.
4.3.6 On confidentiality and the limits of what can be shared
A portion of the play histories that informed Play Styles were taken in clinical settings. These histories are protected by patient confidentiality and cannot be individually identified, shared, or cited. The patterns they contributed are embedded in the framework, but the underlying records are not part of the public domain.
This is not a limitation unique to Play Styles. It is a standard feature of any framework that emerges from clinical practice. The same is true of the foundational work behind many psychological frameworks that are now widely used. The clinical origin of this material is, in fact, a mark of seriousness: these were not casual self-reports but deeply considered personal narratives taken in a context of therapeutic trust.
The adolescent play histories collected in 2020 do not carry the same confidentiality constraints, as they were conducted in a non-clinical peer research context. However, in keeping with our commitment to privacy and research ethics, individual participants are not identified in any published or shared materials.
4.3.7 Additional considerations that strengthen the play history foundation
Several factors reinforce the credibility of the play history foundation:
The play history methodology is consistent with established approaches in narrative psychology and life history research, which have a long track record in the psychological literature.
The patterns Dr. Brown identified have proven durable across time, culture, and context — they appear in play histories taken across five decades, with people from widely varying backgrounds, ages, and life circumstances. The consistency of these patterns across populations as different as clinical psychiatric patients, corporate professionals, Playposium workshop participants, BBC and PBS documentary subjects, and high school students conducting peer research speaks to the robustness of the underlying framework.
The six styles and their subtypes have demonstrated internal coherence in the early assessment data — people who identify strongly with a given style tend to answer style-relevant questions consistently, suggesting the categories are capturing something real rather than arbitrary.
Finally, Play Styles makes no claims beyond what the evidence supports. It is a self-insight tool grounded in decades of careful observation, not a clinical instrument or a predictive assessment. That honesty about scope is itself a form of methodological integrity.
The corpus is qualitative in origin, and the patterns drawn from it are the starting point for the instrument rather than proof of it. Collection is ongoing.
4.3.8 Why play history questions appear in the assessment itself
Because play histories informed this framework so heavily, we chose to incorporate a series of play history questions as part of the assessment design. This is not only to catch a snapshot of a person's current play profile but also to help them reflect on the way their play has evolved across their lifespan. That reflection is what adds the depth and breadth that makes play real for people in a lasting way.
These items are deliberately unscored. They do not touch a style score, a subtype score, a dimension, or the well-being snapshot. They exist to return the instrument to its own origin: the narrative arc that Dr. Brown's play histories surfaced in conversation, offered here in a short structured form that a person can answer on their own. See 7.4.5 for the item set and 7.6 for the design rationale.
4.4 The lineage of personality mapping
The quest to understand human nature is ancient. Hippocrates observed four temperaments in the fifth century BC. In 1921 Carl Jung's Psychological Types (Jung, 1921/1971) proposed that human differences are not random but follow recognizable patterns rooted in how we perceive the world and make decisions, identifying preferences including introversion and extraversion, thinking and feeling, sensing and intuition. For over a century, researchers have tested, refined, and expanded that work into frameworks including MBTI, the Enneagram, the Big Five, DISC, and CliftonStrengths.
Play Styles sits in this lineage but measures something different. Where most instruments describe how a person thinks, decides, or relates, Play Styles describes how a person enters play, and uses that innate, primitive, intrinsically motivated foundation to assess subsequent thinking and behavior patterns.
5. The paradox we hold
Even with centuries of research, personality is dynamic, interconnected, and a mysterious blend of biology and environment, far more complex than any framework can fully capture. This incompleteness has generated substantial debate across disciplines about the scientific grounding of personality instruments generally.
Play Styles provides tools that help us see patterns, find language, and adapt behavior if we choose. This framework will always be incomplete and oversimplified. We offer it with humility, knowing that simplification serves understanding only when we remain aware of the vast complexity beneath.
PART II: THE MODEL
6. The Play Styles model
6.0 From eight play personalities to six Play Styles: an evolution of the framework
Dr. Brown's original published framework, introduced in Play: How It Shapes the Brain, Opens the Imagination, and Invigorates the Soul (2009) identified eight play personality types: the Joker, the Kinesthete, the Explorer, the Competitor, the Director, the Collector, the Artist/Creator, and the Storyteller.
Those eight types represented Dr. Brown's synthesis of the patterns he had observed at that point in his work. They have been widely cited, taught, and applied in the years since publication.
The framework presented in Play Styles reflects a deliberate refinement of that original model. Through continued analysis of play histories, review of the accumulating research literature, and the practical experience of applying the framework across diverse populations and contexts, it became clear that six primary styles with additional subtypes offered a more accurate, workable, nuanced, and comprehensive representation of how people naturally enter the state of play.
Several of the original eight personalities overlapped in ways that created ambiguity in practice; respondents and practitioners sometimes found it difficult to distinguish clearly between them. Consolidating the framework into six primary styles through a robust subtype structure resolves that ambiguity.
Scientific frameworks evolve as evidence accumulates and as real-world application reveals what holds up and what needs refinement. The six Play Styles are research- and experience-informed: they were developed in collaboration with Dr. Brown and shaped by continued analysis of play histories, review of the accumulating research literature, and the practical experience of applying the framework across diverse populations and contexts. They represent a deliberate refinement of the original eight, not a claim that the data alone compels this exact structure.
6.1 The six styles
Each Play Style describes a natural pathway into the state of play, not a static personality type. The definitions below state what the construct is and is not.
Creator — express and make. Creators play by bringing something into being that did not exist before, shaping material, sound, story, image, or a moment of humor into a form that can be shared. Behavioral markers include mentally redesigning ordinary objects, filling blank space, and generating more ideas than a project requires. Under deprivation, Creators report restlessness and irritability, and a sense that everything they do is for someone else. This is not about achievement in a creative profession. The construct is making for its own sake, which is why its structural opposite is Competitor.
Competitor — win and master. Competitors play through contest and the pursuit of mastery, measured against a standard, a clock, an opponent, or a previous best. Behavioral markers include converting ordinary tasks into challenges, keeping score informally, and staying with a problem past the point most people set it down. Under deprivation, Competitors report a gray flatness in which nothing has stakes and effort stops feeling like it counts. This is not competitiveness as a character flaw or an over-emphasis on outcomes. It is a genuine enjoyment of high pressure and challenge where another person might feel stress or fear. The construct is the pleasure of high stakes.
Explorer — discover. Explorers play by encountering the unfamiliar, whether that is a new place, idea, mechanism, feeling, or possibility. Behavioral markers include taking unfamiliar routes, following questions past the point of usefulness, and taking things apart to see how they work. Under deprivation, Explorers report claustrophobia inside their own routine, a sense that every day repeats the last. This is not restlessness or low commitment. The construct is appetite for novelty, which is why its structural opposite is Organizer.
Organizer — create order. Organizers play by arranging, which includes bringing people together, building systems, curating collections, and making plans that work. Behavioral markers include organizing for the pleasure of the arrangement rather than for its outcome, and reaching for structure as a source of satisfaction rather than duty. Under deprivation, Organizers report scattered edges, where the plans and gatherings they usually love begin to feel like chores. This style carries a measurement problem addressed in section 9: organizing is less culturally legible as play, and Organizers rate their own playfulness lowest of any style.
Dreamer — imagine and immerse. Dreamers play through the inner world, using imagination, immersion in other worlds, inhabited characters, and sustained observation. Behavioral markers include elaborate mental scenario-building, deep absorption in narrative, and noticing detail others miss. Under deprivation, Dreamers report a thinning of the inner world, days passing without a single daydream. This is not disengagement or passivity. The construct is interior activity, which is why its structural opposite is Mover.
Mover — move and embody. Movers play through the body, drawn to motion, physical sensation, speed, and the pleasure of being in one's own body with no goal attached. Behavioral markers include moving in order to think, seeking physical sensation, and restlessness when still. Under deprivation, Movers report static in the body, a restlessness that sitting still only sharpens. This is not athletic ability or fitness motivation. The construct is movement for its own sake rather than training toward a target.
6.2 The twenty-four subtypes
- Creator: Joker, Artist, Storyteller, Performer, Musician, Builder
- Explorer: Adventurer, Researcher, Tinkerer, Emotional Explorer, Entrepreneur
- Competitor: Sports Competitor, Gamer, Puzzler, Debater
- Organizer: Social Connector, Director, Coordinator, Collector
- Dreamer: Fantasizer, Role-player, Observer
- Mover: Kinesthete, Thrill-seeker
Subtype scores are normalized within style on a 0 to 10 scale. A person's primary subtype is the top subtype of their primary style; their secondary subtype is the top subtype of their secondary style, not the runner-up within the primary.
Subtype counts are deliberately uneven, from six under Creator to two under Mover, because they follow the granularity observed in the play histories rather than an imposed symmetry. Formal derivation of the subtype structure from response data is listed as outstanding work in section 13.
6.3 The seven dimensions
Solo/Social, Calm/Fast, Receptive/Active, Structure/Free-form, Familiar/New, Physical/Imaginative, Mastery/Exploring. Each is scored 0 to 10.
Six of the seven are asked directly as sliders in Tier 2. The seventh, Solo/Social, is derived from the Tier 1 context slider rather than asked again, so a reader counting items in 7.4.3 will find six.
The dimensions are intended as axes that cut across all six styles, to accurately reveal that two Creators can still play differently from each other. Some clear patterns or overlap with the Play Styles is expected, most obviously between Physical/Imaginative and the Mover/Dreamer distinction. Quantifying that overlap is listed as outstanding work in section 13.
6.4 Structural opposites, derived from data
Every style has a fixed opposite: Creator with Competitor, Explorer with Organizer, Dreamer with Mover.
Unlike most frameworks, these pairings were not assigned theoretically. They were derived empirically from response data and then checked for conceptual coherence. The derivation is reported in full in section 9.1, because it is one of the strongest pieces of evidence in this document that the framework's structure appears in real behavior rather than only in its own design.
6.5 Character strengths
Each style carries three associated character strengths, named using the VIA classification conventions (Peterson & Seligman, 2004): Creator with Creativity, Humor, and Vulnerability; Explorer with Curiosity, Open-mindedness, and Resourcefulness; Competitor with Grit, Ingenuity, and Resilience; Organizer with Teamwork, Communication, and Integrity; Dreamer with Empathy, Imagination, and Self-regulation; Mover with Confidence, Adaptability, and Courage.
These strengths are derived, not measured. They are assigned by style from the theoretical model and the play histories. No item in the instrument measures a respondent's actual level of any of them, and no report should be read as claiming that it does. We state this plainly because the distinction is easy to blur and consequential if blurred.
6.6 A note on model type
The model is best understood as formative rather than reflective (Jarvis, MacKenzie, & Podsakoff, 2003): a style is a composite of related play preferences rather than a latent cause that produces correlated item responses. This matters for how the instrument should be judged. Reflective models are properly evaluated with internal consistency and factor analysis. Formative models are properly evaluated on content coverage and criterion relationships.
We take a middle position. We will report internal consistency and factor structure because they are informative and because readers expect them, while noting that a style scale need not be highly internally consistent to be useful if its items cover genuinely distinct routes into the same kind of play.
PART III: THE INSTRUMENT
7. Codebook and methodology
Generated directly from the production question banks and scoring engine, so item wording and scoring constants here match exactly what respondents see and what is stored in the research dataset.
7.1 Instrument overview
| Free assessment (Tier 1) | Premium assessment (Tier 1 + Tier 2) | |
|---|---|---|
| Purpose | Identify a respondent's dominant Play Style | Identify style profile, subtype profile, play dimensions, and well-being context |
| Items | 18 Likert + 1 attention check (unscored) + 3 forced-choice + 12 context sliders + 1 ranking (6 items) + 4 demographics | All Tier 1 items plus 48 subtype Likert + 6 within-style rankings + 6 dimension sliders + barriers + 4 Play History items (unscored) + 1 optional private reflection |
| Outputs | 6 style scores (0–22), primary style | 6 style scores + ranks, 24 subtype scores (0–10), 7 dimensions, well-being snapshot, Play History narrative section, data-quality flag |
| Administration | Self-report, web, untimed, ~5 min | Self-report, web, untimed, ~15–20 min |
Fixed-form by design. Every respondent receives the identical item set. There is no adaptive branching, no item trimming, and no personalization of the question set. Presentation order is fixed for every block except the Likert blocks (Tier 1 Section 1 and Tier 2 Section A), which are shuffled per respondent using a stable per-session seed so that a refresh preserves the order that respondent saw. Item identity and scoring are unaffected by order.
The six Play Styles
| Key | Name | Theme | Subtypes |
|---|---|---|---|
creator | Creator | Express & Make | Joker, Artist, Storyteller, Performer, Musician, Builder |
explorer | Explorer | Discover | Adventurer, Researcher, Tinkerer, Emotional Explorer, Entrepreneur |
competitor | Competitor | Win & Master | Sports Competitor, Gamer, Puzzler, Debater |
organizer | Organizer | Create Order | Social Connector, Director, Coordinator, Collector |
dreamer | Dreamer | Imagine & Immerse | Fantasizer, Role-player, Observer |
mover | Mover | Move & Embody | Kinesthete, Thrill-seeker |
7.2 Response scales
Likert (Tier 1 Section 1 and Tier 2 Section A) — 5-point, agreement/self-description anchors:
1= Not Like Me at All2= Not Like Me3= Neutral4= Like Me5= Very Much Like Me
Stored as the raw integer 1–5. As of Version 1.1 every Likert item in both tiers is positively keyed, so the scored value equals the raw value. (Waves collected before this revision contain four reverse-keyed Tier 1 items, stored raw and reversed at scoring time as 6 − raw; the raw CSV export for those waves includes both the raw and the reverse-applied value.)
Forced choice — binary A / B, no neutral option, no skip.
Sliders — continuous integer sliders. Tier 1 context sliders are optional and default to the midpoint; a respondent leaving a slider untouched is recorded at its default value, not as missing. Tier 2 dimension sliders run 0–10 between two named poles.
Ranking — drag-to-order, forced complete ordering (no ties, no partial ranks). Position 1 = most appealing.
Demographics — optional, single-select except country (free text, normalized post hoc).
7.2.1 Item parameters at a glance
Every item in the instrument can be described with the same six parameters. Analysts should be able to type a variable straight into a stats package from this table.
| Block | Level of measurement | Range / values | Required | Missing code | Direction | Contribution to score |
|---|---|---|---|---|---|---|
| Tier 1 Section 1 — style Likert | Ordinal, treated as interval | 1–5 integer | Yes (gated) | none possible | All positively keyed | 1–5 pts to one style |
| Tier 1 Section 1 — attention check | Nominal | single correct value | Yes (gated) | none possible | n/a | Unscored; sets data_quality_flag |
| Tier 1 Section 2 — forced choice | Nominal binary | A / B | Yes (gated) | none possible | n/a | +1 pt to the chosen style |
| Tier 1 Section 3 — context sliders | Interval | see 7.3.3, integer | No (fully optional section) | empty string = never rendered; the midpoint is a valid deliberate answer and is not distinguishable from an untouched default | Mixed (burnout, screen_time_erosion are negatively valenced) | Unscored covariate |
| Tier 1 Section 4 — ranking | Ordinal (ipsative) | 1–6, complete order, no ties | Yes (gated) | none possible | 1 = most appealing | 6 pts (rank 1) … 1 pt (rank 6) |
| Tier 1 Section 5 — demographics | Nominal | see 7.3.5 | No | empty string, plus explicit "Prefer not to say" | n/a | Not scored |
| Tier 2 Section A — subtype Likert | Ordinal, treated as interval | 1–5 integer | Yes (gated) | none possible | All positively keyed | 2 items summed → subtype raw 2–10 |
| Tier 2 Section B — within-style ranking | Ordinal (ipsative) | 1–k within each style | Yes (gated) | none possible | 1 = most like me | Weighted into subtype score (7.5.3) |
| Tier 2 Section C — dimension sliders | Interval | 0–10 integer | Yes | none possible | Bipolar, named poles | Dimension score |
| Tier 2 — barriers | Nominal, multi-select | see 7.4.4 | No | empty array | n/a | Not scored |
"Required (gated)" means the UI will not advance until the item is answered, so within a completed submission these variables have no missing values. Partial submissions are never written to the results dataset — abandonment happens before the write, so quiz_results and paid_quiz_results have no partially-completed rows and item-level missingness is structurally zero for gated items. Abandonment is captured separately: a pseudonymous quiz_progress table records the furthest step reached per session for both tiers, and is deliberately held apart from the results dataset so drop-off records can never be mistaken for completed responses. Dropout and attrition analysis should use quiz_progress (surfaced in the admin drop-off view), never the results tables.
7.3 Tier 1 item bank (free + premium)
7.3.1 Section 1 — Style Likert (18 items)
Three items per style (6 styles x 3), presented in an order deliberately scrambled relative to style so respondents cannot infer the construct being measured. All 18 items are positively keyed as of Version 1.1. Each item contributes 1-5 points to exactly one style. A single unscored attention check is embedded in this block and is seen by every respondent in both tiers.
Verbatim item wording, item IDs, keying, and the CSV variable map are withheld from the public edition of this document. The complete item bank is available to researchers, reviewers, and prospective partners under a written agreement: inquiry@nifplay.org.
7.3.2 Section 2 — Forced choice (3 pairs)
Three ipsative A/B pairs, each pitting two styles against each other to break Likert acquiescence. Every style appears in at most one pair, so the maximum forced-choice contribution to any style is 1 point.
Verbatim pair wording is withheld from the public edition.
7.3.3 Section 3 — Context sliders (12 items, unscored)
These do not contribute to style scores. They provide well-being and context covariates for research and drive the premium well-being snapshot. All are optional.
| Variable | CSV column | Range | Default | Low anchor | High anchor | Prompt |
|---|---|---|---|---|---|---|
culture | S3_culture | 0–10 | 5 | Unsupportive | Supportive | The environments I'm part of support and encourage play. |
connection | S3_connection | 0–10 | 5 | Weak | Strong | I have a strong sense of community and belonging. |
engagement | S3_engagement | 0–12 | 6 | 0 hours | 12 hours | In a typical day, how many hours do you feel fully present and engaged in what you're doing? |
burnout | S3_burnout | 0–10 | 5 | Not burned out | Fully burned out | Right now, you feel burned out. |
permission | S3_permission | 0–10 | 5 | Never | Always | You give yourself permission to play as an adult. |
childhood_play | S3_childhood_play | 0–10 | 5 | Not at all | Very much so | Play was a significant part of your childhood. |
playfulness | S3_playfulness | 0–10 | 5 | Not playful | Very playful | You approach your life through a lens of playfulness. |
play_hours | S3_play_hours | 0–20 | 10 | 0 | 20 | In a typical week, how many hours do you spend playing? |
childhood_play_persists | S3_childhood_play_persists | 0–10 | 5 | None of it | All of it | The play you loved earlier in life is still part of your life today. |
tech_effect_on_play | S3_tech_effect_on_play | 0–10 | 5 | Replaced it | Expanded it | Has digital technology replaced or expanded your play? |
screen_time_erosion | S3_screen_time_erosion | 0–10 | 5 | Not at all | Constantly | Scrolling and passive screen time eat into the time you'd otherwise spend playing. |
solo_vs_social | S3_solo_vs_social | 0–10 | 5 | None of it | All of it | How much of your play is done with other people? |
Note the valence: higher is "better" for culture, connection, engagement, permission, childhood_play, playfulness, play_hours, childhood_play_persists, and tech_effect_on_play; higher is "worse" for burnout and screen_time_erosion; solo_vs_social is directionless (a descriptor, not a good/bad continuum). Reverse before building any composite index.
7.3.4 Section 4 — Forced ranking (6 activities)
Respondents order six activities - one per style - from most (1) to least (6) appealing. No ties, no partial ranks.
Verbatim activity wording is withheld from the public edition.
7.3.5 Section 5 — Demographics (optional)
| Variable | Type | Response options |
|---|---|---|
ageRange | Single select | 16–17; 18–24; 25–34; 35–44; 45–54; 55–64; 65+; Prefer not to say |
gender | Single select | Female; Male; Non-binary; Self-describe / Other; Prefer not to say |
country | Free text | Normalized at analysis time (see 7.7.3) |
sector | Single select | Education; Healthcare; Technology; Business / Corporate; Government / Public sector; Nonprofit; Arts / Creative; Hospitality / Service; Trades / Manufacturing; Student; Retired; Other; Prefer not to say |
Minimum eligible age is 16 (the strictest global floor that does not require parental consent). No option below 16 exists.
7.4 Tier 2 item bank (premium only)
7.4.1 Section A — Subtype Likert (48 items)
Two positively-keyed items per subtype across all 24 subtypes. No reverse keying in this section: the items are appetitive preference statements where reversal reads as awkward or double-barreled. The two items for a subtype sum to a raw 2-10.
Verbatim item wording and item IDs are withheld from the public edition.
7.4.2 Section B — Within-style subtype rankings
Six forced rankings, one per style, in which the respondent orders that style's subtypes (2 to 6 of them, depending on the style) from most to least like them. The ranking is a tiebreaker on the Section A Likert items rather than a primary signal.
Verbatim option wording is withheld from the public edition.
7.4.3 Section C — Dimension sliders (6 items)
Six bipolar 0-10 sliders, one per dimension: energy tempo, active/receptive, structure, novelty, engagement mode, and motivation. A seventh dimension, social orientation, is not asked again here - it is derived from the Tier 1 solo-versus-social context slider. The dimension poles are described in section 6.3.
Verbatim slider prompts and pole labels as presented to respondents are withheld from the public edition.
7.4.4 Section E — Barriers (unscored)
Prompt: What most gets in the way of playing more? (Select up to 3)
- Not enough time
- Too tired or low energy
- Feel guilty playing when there's real work to do
- Money
- No one to play with
- Don't know what I'd enjoy
- My work or family culture doesn't support it
- Physical or health limitations
- Other
Multi-select, maximum 3. Descriptive only; never enters scoring.
7.4.5 Section F — Play History (unscored)
Four required items plus one optional open reflection, presented as their own section of the premium assessment. Nothing in this block enters any score, rank, dimension, or well-being value, and no item order or item set depends on the answers. Selections drive which pre-written narrative blocks appear in the "Where your play comes from" section of the report; the mapping is a fixed lookup, not a scoring model.
F1 — Your play story (single select, required). Which of these sounds most like your play story?
| Key | Option |
|---|---|
continuous | I have always played, and I still do. |
interrupted | Something stopped my play, and it has not really come back. |
dormant | Play faded out slowly. I could not tell you exactly when. |
rekindling | There was a long gap, but play is finding its way back in. |
redirected | I kept doing the things I loved, but they turned into work. |
late_start | I did not get to play much as a kid. |
F2 — Childhood play (multi-select, required). As a kid, what did your play usually look like? Options map to the six styles — making, competing, exploring, organizing, imagining, moving — plus "I did not get much time to play" and a free-text "Something else." Up to three selections render in the report, in display order. When a selection matches the respondent's dominant style, a continuity-framed block is shown instead of the neutral one; this is copy selection only and does not alter the dominant style, which is already fixed by the Tier 1 and Tier 2 scoring.
F3 — What shaped your play (matrix, required). What has shaped your play over the years? Twelve rows — school, career, kids, a romantic partner, friends or community, caregiving, health, loss, a big move, when play got competitive, time/money/space, and being told play was not for them — each optionally marked opened up play or shut down play. An exclusive "Nothing has changed the way I play" option and a free-text "Something else" are also available. Up to two blocks per direction render in the report, in row order.
F4 — Who was there (single select, required). Growing up, who was usually there when you played? Mostly other kids; mostly alone and preferred it; mostly alone by circumstance; mostly adults or older siblings; a mix; or did not get to play much.
Open reflection (optional, private). A free-text prompt inviting the earliest remembered play memory, capped at 5,000 characters. It is stored in a separate owner-only table, is excluded from the research export and from every admin and partner view, is never used for research, and is hard-deleted with the account. It appears only at the bottom of the respondent's own report.
7.5 Scoring logic
7.5.1 Tier 1 style score (0-22 per style)
Each style receives a composite of three contributions: its Likert items (all positively keyed as of Version 1.1), its position in the Section 4 ranking, and any forced-choice pick that credits it. Theoretical range is 4-22. The reported percentage is the score expressed as a share of the maximum and is used only for bar widths - it is not a normative percentile.
Weighting rationale: self-description (Likert) carries the large majority of the maximum; the ipsative items act as a corrective on flat or acquiescent response sets rather than as a primary signal.
Exact formulas, point weights, and constants are withheld from the public edition.
7.5.2 Style ranking and tiebreak cascade (premium)
Styles are ordered into a unique 1-6 rank using a deterministic cascade: composite score first, then the respondent's own Section 4 ranking, then forced-choice confirmation, then a fixed canonical order as a final fallback. Primary style = rank 1, secondary = rank 2, "biggest stretch" = rank 6. No random tiebreaking is used anywhere, so re-scoring an archived answer set always reproduces the original report.
7.5.3 Subtype score (0-10, normalized within style)
A subtype's raw total combines its two Section A Likert items with its position in the within-style ranking, then is normalized to a 0-10 score with one decimal. Because styles contain different numbers of subtypes (2 to 6), raw totals are not comparable across styles; the within-style normalization makes the report bars comparable. The Likert contribution dominates by construction; the ranking exists mainly to break ties among subtypes a respondent rated identically.
Primary subtype = highest-scoring subtype within the primary style. Secondary subtype = highest-scoring subtype within the secondary style (not the second-highest overall), so the pair describes two different territories.
Exact formulas and normalization constants are withheld from the public edition.
7.5.4 Dimensions
Each 0–10 slider is banded: 0–3 = strongly Pole A, 4–6 = balanced, 7–10 = strongly Pole B. "Balanced" is treated as a substantive finding, not missing data. A respondent whose seven dimensions are all balanced is flagged as a flat profile and receives all-rounder interpretation copy instead of pole-specific copy.
7.5.5 Well-being snapshot
Derived entirely from the Tier 1 context sliders, normalized to a common 0-10 scale. Respondents are placed into one play band (covering, among others, thriving, growing, low-hours, and long-gap patterns) and one well-being band. Band assignment uses strict priority rather than a composite index, so a single acute signal - especially burnout - is never averaged away by otherwise healthy scores.
Exact band definitions and thresholds are withheld from the public edition.
7.5.6 Data quality flag
Both tiers contain one embedded attention check. Records that fail it are retained, scored, and shown to the respondent as normal; the flag exists so analysts can filter or sensitivity-test. Recommended practice: report primary analyses both with and without flagged cases.
The location and correct response of the attention check are withheld from the public edition.
7.5.7 Play History (no scoring)
Section F contributes nothing to any computed value. There is no play-history score, index, or band weighting, and the responses do not modify style scores, subtype scores, dimensions, the well-being snapshot, or the data-quality flag. Report output is a deterministic lookup from the selected options to fixed copy blocks: one block for F1, up to three for F2, up to two per direction for F3, and one for F4. The only conditional in the whole block is F2, where a childhood selection matching the already-determined dominant style swaps in continuity-framed wording. The optional open reflection is rendered verbatim to its author and to no one else.
7.6 Design rationale
Why Likert is the baseline scale. The 5-point Likert (Likert, 1932) is the most extensively validated self-report format in psychological measurement: it is intuitive for respondents with no training, produces roughly interval-level data suitable for correlation, factor analysis, and reliability estimation, and yields normative scores that can be compared across people and across time. A labeled 5-point scale with a true neutral midpoint ("Not Like Me at All" → "Very Much Like Me") is short enough to avoid discrimination fatigue and long enough to capture graded intensity, which is why it anchors both tiers of this instrument.
Summarized for the general reader in section 2 under "A combination of scales." This section gives the analyst-facing detail.
Why forced choice and sliders sit on top of the Likert. Each format captures a different mode of thought, and each fails in a different way, so triangulating across them is more robust than deepening any one of them. Likert items ask how much does this describe me (absolute self-report, vulnerable to acquiescence and flat profiles). Forced-choice A/B pairs ask which of these two is more me — a trade-off judgment that requires the respondent to spend a preference they cannot give to both options, which breaks ties Likert alone leaves unresolved. Rankings extend that logic across a full set. Continuous sliders capture magnitude on context and state variables where a 5-point scale would be too coarse and where no "right" answer exists to agree with. Together, the absolute, the comparative, and the continuous converge on a profile that any single format would render less sharply.
Why an embedded attention check. One instructed-response item is embedded in the Tier 1 Likert block — so it covers free and premium respondents alike — with a single correct answer. It costs the respondent seconds, is invisible to anyone reading carefully, and gives analysts an objective, pre-registered basis for identifying careless or automated responding rather than relying on post hoc judgment (Oppenheimer, Meyvis, & Davidenko, 2009). Combined with completion-time and straight-lining diagnostics, it lets us quantify data quality in every wave and sensitivity-test every published result with and without flagged cases (see 7.5.6).
Why every item is now positively framed. Earlier versions measured each style with a mix of positively worded statements and at least one negatively framed, reverse-scored statement, on the standard rationale that reversals reduce acquiescence bias and interrupt response-set momentum. The August 2026 psychometric evaluation (section 10) showed that rationale did not hold here: two of the four reverse-keyed items loaded on each other rather than on their intended traits, forming a wording artifact that consumed a whole factor. Reversals were also the items respondents most often misread. In Version 1.1 all reverse-keyed items were rewritten as positively keyed statements of the same underlying disposition, and no Likert item in either tier is reverse scored. Acquiescence is instead addressed by the ipsative blocks (forced choice and ranking) and by straight-lining diagnostics.
Why items are short, plain, and non-compound. Every item is written as a single observable idea in plain language, with no double-barreled clauses ("I like building and organizing"), no negations stacked on negations, no jargon, and no clause that requires holding two conditions in mind at once. Compound items are unanswerable when the two halves are true to different degrees, and the resulting variance is uninterpretable — the analyst cannot tell which half the respondent answered. Plain, single-idea wording also improves comprehension for non-native English speakers, keeps reading level low, and is a precondition for the item-level discriminant validity analysis in 7.7.2, which assumes each item indexes one construct. It also translates cleanly. We have added a translation plug-in so the instrument can be read in other languages as accurately as possible, and short, literal, single-idea items are what make that translation reliable.
Why a mixed-format instrument. Pure Likert instruments are vulnerable to acquiescence and to flat profiles where every style scores near-identically. Pure ipsative instruments (rankings, forced choice) produce scores that cannot be compared across people. Combining them keeps the score normative and interpretable while using the ipsative blocks to force differentiation.
Why scrambled presentation order. Items are interleaved across styles so respondents cannot detect the six-block structure and self-curate an identity. This costs nothing psychometrically and materially reduces socially-desirable clustering.
Why subtype names are hidden in Section B. Labels such as "Director" or "Collector" carry connotation. Ranking the behavior instead of the label keeps the item measuring preference rather than label affinity.
Why unscored context sliders. Burnout, connection, permission to play, screen time, and childhood-play continuity are the research questions the Institute cares about most, but they are not style traits. Keeping them out of the style score prevents contaminating a trait measure with a state measure while preserving them as covariates.
Why a Play History block sits inside a trait instrument. The framework itself came out of play histories, so an instrument that measured only present-day preference would return less than the method it was built on. The four Play History items reconstruct the lifespan arc — the shape of the story, what childhood play looked like, what opened it up or shut it down, and who was there — in a form short enough to answer without an interviewer. They are unscored by design: they carry no weight in any style, subtype, dimension, or well-being output, and no branching depends on them. Their value is interpretive and reflective rather than psychometric, which is also why the block renders as narrative copy in the report rather than as a score. The optional open reflection that closes the block is private to the respondent, is never used for research, and never appears in any team or partner output.
Why the report is descriptive, not diagnostic. The instrument is a preference profile, not a clinical or employment-screening tool. No score is presented as a percentile against a norm group, and no language implies pathology. Well-being items are explicitly framed as context, and the report carries a non-clinical disclaimer.
7.7 Validation plan
This section defines the analyses the raw export is designed to support. The Tier 1 reliability, factor-structure, discriminant, convergent, and criterion analyses described here have now been carried out and are reported in section 10; the plan is retained because it also governs the Tier 2 subtype items, which have not yet been evaluated at volume, and the wave 2 re-validation.
7.7.1 Reliability
- Internal consistency. Cronbach's α and McDonald's ω per style (3 Tier 1 items) and per subtype (2 Section A items). With only 2–3 items per scale, expect modest α; report inter-item correlation and mean inter-item correlation alongside it, which is the more defensible statistic for short scales.
- Test–retest. Retakes are tracked (
retake_number,is_retake,previous_result_id), enabling stability coefficients for style scores and primary-style agreement (Cohen's κ) across occasions.
7.7.2 Validity
- Discriminant validity. Item-level correlation matrix across all 18 Tier 1 items and all 48 Section A items. Each item should correlate more strongly with its own scale (corrected item–total) than with any other scale. The raw CSV exports every item as its own column with its question text attached specifically to make this a one-step analysis.
- Convergent validity. Section A subtype items should correlate with their parent style's Tier 1 items; Section B ranking positions should correlate with Section A Likert sums for the same subtype.
- Structural validity. Exploratory factor analysis / parallel analysis on the Tier 1 item pool to test whether six factors emerge, followed by confirmatory factor analysis on a holdout sample. Subtype structure tested within style.
- Criterion / concurrent validity. Correlate style and dimension scores with the context sliders (playfulness, permission, burnout, connection) and, where available, with Live Playfully Journal behavior — daily play frequency, energy change, and self-reported play type. The journal is the strongest behavioral criterion currently available.
- Known-groups. Compare style distributions by sector, age band, and partner cohort.
7.7.3 Data preparation conventions
- Country is free text and is normalized to a canonical country at analysis time; city and region entries (e.g. "Los Angeles") are mapped to their country.
- Duplicate respondents are identified by email (case-insensitive); retake chains are linked through
previous_result_id. - Flagged cases (
data_quality_flag) and anonymized cases (anonymized_atset) should be handled explicitly in every analysis. - Missing sliders: an untouched optional slider is stored at its midpoint default and is indistinguishable from an intentional midpoint. Treat midpoint slider values with appropriate caution in any analysis that hinges on them.
7.8 Dataset and variable naming
Sections 7.8 and 7.9 are written for an analyst working directly with the raw export. Other readers can move on to Part IV without losing the thread.
The admin full raw CSV export flattens every item into paired question/response columns, with the scored value alongside the raw response wherever reverse keying applies. Identity, provenance, and consent fields (participant ID, tier, source partner, retake number, consent timestamp, anonymization timestamp, submission time) lead the file, followed by the six composite style scores and the untouched source payloads. Question text is repeated on every row deliberately: the export is meant to be fully self-describing without a separate key.
Conventions: UTF-8 CSV with a header row, comma-delimited, quoted fields; timestamps are UTC ISO 8601; booleans are true/false; missing is an empty string (never NA, 0, or -99); one row = one completed submission, not one row per person. Tier should always be stratified or controlled for, because premium respondents are self-selected and paid.
An anonymized export (no name or email, sequential respondent IDs) is available for sharing outside the Institute.
The full column dictionary and variable-naming scheme are withheld from the public edition and are supplied with any approved data share.
7.9 Analyst handoff package
Any approved analysis is supplied with the regenerated codebook, the anonymized raw CSV, the date range and filters applied, respondent counts by tier and data-quality flag, and a stated first research question.
Sample size guidance. Item-level reliability and correlation work is usable from roughly n = 200; exploratory factor analysis on the Tier 1 items targets n >= 400; confirmatory factor analysis on a holdout targets n >= 800 total; subtype structure needs the same per-item ratio but draws only on premium respondents, so it lags considerably; known-groups comparisons need n >= 30 per cell.
Known limitations we state up front.
- No external norm group. Scores are interpreted within the respondent, not against a standardized population. There are no percentiles or cut scores.
- Short scales. Three Likert items per Tier 1 style and two per subtype. Expect modest alpha; report mean inter-item correlation alongside it.
- Self-selected, non-probability sample. Respondents arrive through the Institute's channels and partner links; the sample is not representative of any general population, and premium respondents are additionally selected on willingness to pay.
- Mixed item formats. The composite mixes normative (Likert) and ipsative (ranking, forced-choice) contributions, which violates the independence assumptions of some classical statistics. Analysts wanting a clean psychometric model should analyze the Likert block on its own and treat the ipsative blocks as separate evidence.
- Slider midpoints are ambiguous for the optional Tier 1 sliders (see 7.7.3).
- English is the language of record. A translation plug-in makes the instrument readable in other languages, but every item is written and scored in English, and no cross-cultural equivalence testing has been done.
- No convergent measure by design. The instrument is deliberately not administered alongside a clinical or established playfulness scale. Play Styles is a self-insight and language tool, not a clinical or diagnostic instrument, so the goal is internal coherence, reliability, and practical usefulness rather than convergence with a validated psychological construct.
Deliverables worth requesting: reliability table by scale; corrected item-total correlations; full item correlation matrix; EFA scree/parallel analysis with rotated loadings; CFA fit indices on a holdout; correlations of style scores with the context sliders; known-groups comparisons; and an explicit list of items recommended for revision or removal with the reasoning.
7.10 Ethics and data protection
Participation is voluntary and self-initiated; consent timestamps are stored per response. Minimum age is 16. Identified free and premium assessment responses are automatically anonymized after five years, and respondents can export or delete their data self-serve from their account. Well-being items are not diagnostic and are never presented as clinical findings. Data is not sold. Aggregate and de-identified data may be used for research and publication; partner-sourced responses are tagged with a partner code and also flow into the master dataset. Identifiers are stripped after five years, implemented through the anonymized_at field, while the de-identified scores are retained indefinitely. That single decision serves both commitments at once: it is what limits how long we hold personal information, and it is what makes long-run research on the dataset possible.
PART IV: EVIDENCE
8. Pre-launch simulation
Before launch, the complete scoring pipeline was run against 2,000 synthetic respondents with randomized response profiles. The simulation confirmed that:
- The scoring cascade produced a deterministic primary and secondary style for every possible response pattern, including full-tie patterns.
- The tiebreak cascade resolved ties consistently and in documented priority order.
- Subtype normalization behaved correctly at both extremes of the scale.
Stated limitation. The simulation validates scoring machinery, not construct validity. Synthetic respondents cannot confirm that items measure real psychological differences. That evidence requires real respondents, which the next section begins to provide.
9. Preliminary empirical analysis
Sample. N = 3,469 completed free assessments, export dated 13 July 2026. The export contains full style rank orders (1 through 6), wellbeing self-reports (burnout, connection, engagement, playfulness, each 0 to 10), weekly play hours, childhood play significance, and optional demographics. Raw composite style scores were not included in this export, so all style-structure analyses below are rank-based.
Status of this section. Sections 9.1 and 9.2 were re-run in September 2026 on the full sample of 20,966 responses, using composite scores rather than ranks. That re-analysis is reported in section 9.6 and supersedes the figures below where the two differ. The pairing derivation is confirmed; the opposite co-lead rate in 9.2 is corrected there. Sections 9.3 through 9.5 — play and wellbeing, style-level wellbeing differences, and demographic patterns — appear in the internal edition of this document; they are withheld from the public edition pending a larger research report.
Analytic note on ranks. With six forced ranks, any two styles carry a built-in negative correlation of approximately r = -0.20 from the compositional constraint alone. Opposition therefore means correlations meaningfully below -0.20. Correlations near zero indicate affinity, not independence. Every figure below is interpreted against that baseline.
9.1 Derivation of opposite style pairs
Two independent rank-based methods were used.
Method 1: Spearman rank correlations. The most opposed pairs, beyond the -0.20 baseline, were Creator-Mover (-0.30), Explorer-Organizer (-0.27), Dreamer-Mover (-0.27), Creator-Organizer (-0.25), and Creator-Competitor (-0.22). The most affiliative pairs were Competitor-Organizer (-0.10), Competitor-Dreamer (-0.10), Explorer-Mover (-0.10), and Creator-Explorer (-0.12).
Method 2: Optimal symmetric assignment. The product requires three symmetric pairs, so that every style has exactly one opposite. All 15 possible three-pair partitions of the six styles were evaluated for total opposition. The optimum:
Creator ↔ Competitor, Dreamer ↔ Mover, Explorer ↔ Organizer (summed r ≈ -0.75)
Corroboration through lift-normalized last-place rates. For each primary style, the probability of every other style ranking last was computed and normalized against that style's base rate of ranking last. All three selected pairs show mutual opposition. Mover primaries rank Dreamer last at 1.93 times base rate, and Dreamer primaries return 1.32. Organizer primaries rank Explorer last at 1.76, returned at 1.45. Competitor primaries rank Creator last at 1.69, returned at 1.20.
An honest caveat. The single strongest directional repulsion in the data is Competitor primaries ranking Explorer last, at 2.52 times base rate. That pairing was not selected, for two reasons. It is one-directional, since Explorer primaries do not reciprocate. And Explorer's own mutual opposite is clearly Organizer. The symmetric assignment prioritizes mutuality, which the relational use case requires. We report the discarded finding because a reader is entitled to see what the optimization traded away.
Conceptual coherence. The derived pairs map cleanly onto the framework's dimensional structure: Creator against Competitor is making for its own sake against playing to win; Dreamer against Mover is imaginative against embodied engagement; Explorer against Organizer is discovery against order. The empirical result and the conceptual model agree, which is the outcome we hoped for and not one we could have guaranteed.
9.2 Profile combination prevalence
Across the 30 ordered primary-secondary combinations, prevalence ranges from Creator-Explorer at 7.8% to Competitor-Explorer at 0.61%, roughly a thirteenfold spread.
Prevalence is direction-sensitive. Creator-Competitor occurs in 2.42% of profiles, about 1 in 41, while Competitor-Creator occurs in 0.84%, about 1 in 119. The ordered pair matters, and the reporting engine matches on primary-then-secondary order accordingly.
Structural validation. Opposite styles co-lead less often than chance. The rarest combinations in the dataset are precisely the opposite pairs and near-opposites. The framework's internal structure is visible in naturally occurring response patterns, which is a meaningful check on a model that could otherwise be accused of only reproducing its own assumptions. (This section originally put the opposite co-lead rate at under 3%. On the full sample it is 16.0%; see the correction in section 9.6.)
9.6 Full-sample re-analysis (September 2026, N = 20,966)
Sample. Every completed free assessment held on 3 September 2026: N = 20,966, six times the July export. This analysis uses raw composite style scores, not ranks alone, closing limitation 2 in section 12. Two figures are reported per pair: the ipsatized correlation (each respondent's six scores centered on their own mean, removing elevation and acquiescence) and the rank correlation (Spearman, directly comparable to the July method).
Opposition, all 15 pairs.
| Pair | Ipsatized r | Rank r | Framework pair |
|---|---|---|---|
| Creator – Mover | -0.33 | -0.31 | |
| Dreamer – Mover | -0.32 | -0.31 | opposite |
| Explorer – Organizer | -0.31 | -0.25 | opposite |
| Creator – Organizer | -0.28 | -0.26 | |
| Explorer – Competitor | -0.27 | -0.25 | |
| Competitor – Mover | -0.22 | -0.21 | |
| Creator – Competitor | -0.22 | -0.21 | opposite |
| Creator – Dreamer | -0.22 | -0.20 | |
| Organizer – Mover | -0.19 | -0.18 | |
| Organizer – Dreamer | -0.17 | -0.17 | |
| Explorer – Dreamer | -0.14 | -0.20 | |
| Creator – Explorer | -0.12 | -0.15 | |
| Competitor – Dreamer | -0.07 | -0.04 | |
| Competitor – Organizer | -0.06 | -0.12 | |
| Explorer – Mover | -0.05 | -0.08 |
At this sample size every coefficient is statistically distinguishable from zero; magnitude, not significance, is what carries information.
The pairing map holds. Re-running the optimal symmetric assignment over all 15 partitions, the shipped map sums to -0.85 on ipsatized scores against a mathematical optimum of -0.86 (-0.76 against -0.81 on ranks). Dreamer ↔ Mover and Explorer ↔ Organizer are two of the three most opposed pairs in the data. The July derivation is confirmed on a sample six times larger and on the score metric that export could not supply.
What the optimizer would trade, and why we do not. Both optima would pair Creator with Organizer and Competitor with Explorer — the same tension section 9.1 declined, and smaller now: the strongest one-directional repulsion (Competitor primaries ranking Explorer last) has fallen from 2.52 to 2.05 times base rate. Creator ↔ Competitor remains the weakest shipped pair (-0.22, seventh of 15) but is mutual, conceptually clean, and the pair the report's copy is built on. We keep it and state the cost: a purely data-driven map would pair Creator with Organizer.
Prevalence across the 30 ordered combinations runs from Dreamer-Explorer at 7.58% (1 in 13) to Competitor-Explorer at 0.72% (1 in 139), a tenfold spread. Direction sensitivity persists and is sharper: Creator-Competitor 1.88% (1 in 53) against Competitor-Creator 0.76% (1 in 132). Under the shipped banding, 9 combinations are common, 14 uncommon, 7 rare; the smallest cell holds 151 respondents, so the 30-observation guard now passes everywhere. Primary-style prevalence: Dreamer 27.5%, Creator 22.6%, Explorer 15.9%, Mover 15.4%, Organizer 14.0%, Competitor 4.7%.
Correction to the July co-lead figure. Section 9.2 put opposite pairs co-leading a profile at under 3%. On the full sample the figure is 16.0% — about 1 in 6, against a 20% random baseline (Dreamer + Mover 7.62%, Explorer + Organizer 5.72%, Creator + Competitor 2.64%). The direction survives — opposites co-lead less than chance — but the magnitude does not: carrying your own opposite is uncommon, not rare, and the report's "you carry your own opposite" variant fires for roughly one reader in six.
Dimension-to-style correlations. Only one of the seven dimensions, Social Orientation, is carried on the Tier 1 form; the other six are Tier 2 items with too few completions to analyze. The table below therefore reports Social Orientation together with the Tier 1 context sliders, against ipsatized style scores (n = 17,537 to 20,966 depending on slider coverage). Cells at ±0.10 or above:
| Slider | Association |
|---|---|
| Social orientation (solo → social) | Mover +0.15, Dreamer -0.14 |
| Playfulness | Organizer -0.18, Explorer +0.16, Creator +0.11 |
| Engagement | Dreamer -0.20 |
| Connection | Dreamer -0.14, Organizer +0.11 |
| Burnout | Dreamer +0.12 |
| Screen-time erosion of play | Dreamer +0.14 |
| Technology's effect on play | Competitor +0.14, Mover -0.10 |
| Permission to play | Organizer -0.12, Explorer +0.09 |
| Weekly play hours | Organizer -0.13, Competitor +0.10 |
| Childhood play persists | Creator +0.12, Organizer -0.12 |
Every coefficient is small — the largest accounts for four percent of variance — which answers the question posed in section 13: the dimensions and context sliders are not redundant with the styles. They measure something adjacent, as the two-layer model requires. The table also extends the July style-level wellbeing findings: the Organizer playfulness deficit, the strongest style-slider association in the instrument, now appears alongside lower permission to play and fewer weekly play hours — supporting the permission-and-legibility reading over a deficit. The Dreamer pattern (higher burnout, lower engagement and connection, more screen-time erosion, more solo) is consistent across five independent sliders, a real profile rather than a single noisy finding. The Tier 2 dimensions remain unevaluated (section 13).
Limits of this re-analysis. Same convenience sample, same self-report method, same single sitting. The composite scores are still partly compositional, which is why ipsatized and raw figures are reported separately. Nothing here is test-retest evidence, and nothing here re-opens the Version 1.0 Competitor scale problem documented in section 10.2; Competitor findings in this section inherit that provisional status.
10. Psychometric evaluation of the Tier 1 item bank (August 2026)
This is the first full psychometric evaluation of the instrument on real respondents. It was run on 4,127 free-tier submissions collected 26 July to 12 August 2026 and reduced to a cleaned analytic sample of N = 3,730 after removing 9 test accounts and 388 repeat submissions (earliest attempt kept). There was no item-level missingness, because every scored item is gated in the interface. The sample was split at random into an EFA half (n = 1,865) and a CFA holdout half (n = 1,865).
The analysis followed the validation plan in section 7.7: reliability, exploratory factor analysis, confirmatory factor analysis on a holdout, convergent and discriminant validity, graded-response IRT, and criterion validity against the context sliders using Johnson's Relative Weights Analysis with 2,000 bootstrap resamples. The Likert block was treated as the psychometric core; the ranking and forced-choice blocks were held out of the factor models and used only as convergent evidence, because mixing normative and ipsative components violates the independence assumptions those models rest on.
What has already been fixed. This evaluation was run on the Version 1.0 form of the instrument. Two of its findings have since been acted on: all reverse-keyed items were removed, and the Competitor scale was rewritten. Those revisions are live and are documented in section 10.8. Everything reported in 10.1 through 10.7 describes the instrument as respondents saw it in wave 1, not as it stands today.
Scoring engine verified. Every stored composite score was reproduced exactly from the raw answers, and all Likert responses fell within the valid 1–5 range. Whatever else this evaluation found, the scoring pipeline itself is doing what the codebook says it does.
10.1 Headline findings
- Five of the six styles are real and separable. Creator, Explorer, Organizer, Dreamer, and Mover each emerge as a distinct factor in the exploration half, replicate in the confirmatory holdout, and show clean discriminant validity — every HTMT ratio is far below the .85 threshold at which two constructs are considered empirically indistinguishable.
- The Competitor scale did not work in its Version 1.0 form. This is the single most important finding in the document. The scale has since been rewritten (section 10.8) and is being re-validated on wave 2.
- Reliability of the other five scales is modest but normal for three-item scales, and the fix is length rather than content.
- Overall model fit is below conventional standards (CFI = .79 in the holdout), driven almost entirely by Competitor and five specific weak items. Three of those items, and both problem reverse items, have since been rewritten.
- Scores relate to real-world context in sensible, style-specific ways — different styles predict different outcomes, which is itself validity evidence.
10.2 Reliability
| Scale | α | Ordinal α | Mean inter-item r |
|---|---|---|---|
| Creator | .64 | .69 | .37 |
| Explorer | .51–.58 band | .57–.69 band | .26–.33 |
| Organizer | .51–.58 band | .57–.69 band | .26–.33 |
| Dreamer | .51–.58 band | .57–.69 band | .26–.33 |
| Mover | .51–.58 band | .57–.69 band | .26–.33 |
| Competitor | .26 | — | .11 |
Creator is the strongest scale. Explorer, Organizer, Dreamer, and Mover sit in the modest-but-workable range expected of three items (α = .51 to .58, mean inter-item r = .26 to .33). None reaches the conventional .70 comfort level, but by the Spearman-Brown projection, extending those scales to five items of similar quality would lift α to roughly .69 to .75. That is a length problem, not a coherence problem.
Competitor is a different case. In its Version 1.0 form, an α of .26 with a mean inter-item correlation of .11 means its three items barely relate to one another: a respondent who endorses one is almost no more likely to endorse the others. Lengthening cannot repair a scale whose items do not share a construct, which is why the scale was rewritten rather than extended (section 10.8).
10.3 Factor structure
Horn's parallel analysis retains five factors, not six, identically on polychoric and Pearson correlations. The six-factor solution makes the reason concrete: five clean factors corresponding to Creator, Explorer, Organizer, Dreamer, and Mover, plus a sixth factor that is not Competitor at all but a reverse-wording doublet — two reverse-keyed items clustering with each other rather than with their intended traits. All reverse-keyed items have since been removed from the instrument (section 10.8).
Confirmatory fit on the holdout: χ²(120) = 1,021, CFI = .787, TLI = .729, RMSEA = .063, SRMR = .058. The residual-based indices (RMSEA, SRMR) indicate the five healthy scales reproduce their correlations reasonably well; CFI and TLI are harsher judges here because they benchmark against a zero-correlation baseline that short, deliberately heterogeneous scales are always going to sit close to. We report both rather than the flattering one.
Convergent validity per factor: AVE = .33 to .40 and composite reliability = .56 to .66 for the five healthy scales — below the classic .50 AVE comfort threshold, as expected for three short items, but serviceable. Competitor fails outright (AVE = .14, CR = .26).
Discriminant validity is the clearest good news in the evaluation: every HTMT ratio falls between −.01 and .53, far below .85. Whatever the scales measure, they are not measuring the same thing as each other.
10.4 Item-level diagnostics
Graded-response IRT (Samejima) was fitted to the Likert items as an independent check on the factor analysis, and the two traditions agree. Each healthy scale rests on two strong anchors (discrimination a = 1.4 to 2.4, loadings .54 to .86) plus one weaker third item. All three Competitor items discriminate poorly (a = .53 to .70) — its best item is weaker than every other scale's weakest.
Five items were flagged as weak across all three methods (loading, IRT discrimination, item-total correlation). Three of them — both Competitor items and the Mover reverse item — have since been rewritten (section 10.8); the Dreamer and Explorer items below are unchanged and remain under watch:
| Item | Scale | λ | IRT a | Diagnosis |
|---|---|---|---|---|
| Competitor reverse item | Competitor | .14 | .68 | Lowest sampling adequacy in the battery; clusters with the reverse-wording doublet. It measures emotional regulation after losing, not competitiveness. |
| Mover reverse item | Mover | .30 | .64 | Forms a pseudo-factor with the Competitor reverse item. Sitting tolerance is not the opposite of embodied play. |
| Competitor challenge item | Competitor | .22 | .53 | Challenge-as-novelty wording pulls it toward Explorer. |
| Dreamer absorption item | Dreamer | .29 | .76 | Near-universally endorsed, so it separates almost nobody, and it relates only weakly to the imaginative core. |
| Explorer restlessness item | Explorer | .35 | .77 | Restlessness wording leaks onto the reverse-wording doublet. |
Two further items are keep-but-monitor: one Creator item cross-loads on Dreamer (ideating versus making), and one Competitor item behaves as conscientious self-tracking, correlating more with Organizer than with its own scale.
Two of the four reverse-keyed items behaved as a wording artifact rather than as trait measures. The other two reverse items perform well (loadings .66 and .54). The difference is instructive: the clean reversals negate the trait directly, while the failed reversals negate a behavior merely adjacent to it. On the strength of this finding all four reverse items were replaced with positively keyed statements in Version 1.1, so no Likert item in either tier is now reverse scored.
A thresholds problem that affects every scale. Most items are very easy to endorse — first thresholds frequently below −3. Measurement information is concentrated among low-to-average scorers, so the instrument discriminates least well exactly where users care most: at the top, among people deciding which style is truly theirs. Future items need more demanding wording (behavioral frequency, comparative claims) to extend information upward.
10.5 Convergence across formats
The Likert, ranking, and forced-choice blocks agree moderately and in a style-dependent way: strongest for Mover (.51) and Creator (.45), weakest for Competitor (.19). Forced choices match the corresponding Likert difference 64% to 71% of the time. Two readings are compatible and both are probably true: the formats capture overlapping but non-identical information, which is the design intent, and single-item ipsative blocks are noisy. Either way, the pattern supports the codebook's decision to let the Likert block dominate the composite.
10.6 Criterion validity
Relative Weights Analysis partitions each outcome's explained variance among the six correlated styles.
- Playfulness and permission to play are the best-predicted outcomes (R² = 9.5% and 3.9%), and the credit is concentrated: Explorer carries 47% of the playfulness prediction (95% CI 37–56%) and Creator another 27% (CI 18–35%).
- Community connection is carried by a completely different pair: Mover (40%) and Organizer (39%) split nearly all of it.
- Burnout, screen-time erosion, and lower daily engagement are driven by Dreamer alone, which accounts for 79% of the (small, R² = 2.5%) burnout prediction. Higher Dreamer goes with more burnout (r = +.14), more screen-time erosion (+.17), and fewer engaged hours (−.13).
- Competitor predicted nothing in wave 1, consistent with its measurement failure. A scale with no internal coherence cannot show external validity.
That crossover pattern — different styles owning different outcomes — is itself validity evidence: the six scores are not interchangeable. The absolute effect sizes are modest (R² of 2.5% to 9.5%), which is the expected ceiling for three-item trait scales predicting single-item state sliders, both sides attenuated by measurement error. It should improve mechanically as reliability improves.
The Dreamer finding needs care, not celebration. The scale's strongest marker is mind-wandering, and unmanaged mind-wandering is a known correlate of low engagement. The scale may be blending playful imagination with attentional drift. This is now the leading candidate for an item-level split, so that the style's report copy never inadvertently celebrates a disengagement pattern.
10.7 Limitations specific to this evaluation
- Self-selected free-tier sample, 63% female, reached through the Institute's own channels. All conclusions are about the instrument's internal behavior in this population, not about population prevalence of styles.
- All criteria are concurrent self-reports collected in the same sitting. Common-method variance inflates, and single-item criterion unreliability deflates, the criterion correlations. The RWA percentages, being relative shares, are more trustworthy than the absolute R² values.
- De-duplication by name may merge distinct people who share a name; results were confirmed stable under an email-only rule.
- Three-item scales cap attainable reliability and make the IRT calibrations approximate. The GRM parameters are diagnostics, not a calibrated item bank.
- No external convergent instrument is administered, by design, so validity here is internal, structural, and criterion-based within the assessment's own ecosystem. Test-retest stability could not be assessed at all, because retake linkage was broken.
- Tier 1 only. The premium Tier 2 subtype items have not yet been evaluated at volume.
A second evaluation is planned. The item revisions in section 10.8 have now been implemented and a second wave of data collection is under way, after which this section will be extended with a revised-instrument evaluation reported the same way, whatever it shows.
10.8 Revisions implemented (Version 1.1, August 2026)
The following changes are live in the instrument. They are reported here so that any analysis can be tied to the exact form a respondent saw.
- All reverse-keyed items removed. The four reverse-keyed Tier 1 items - one each for Mover, Creator, Competitor, and Organizer - were rewritten as positively keyed statements of the same disposition. No Likert item in either tier is now reverse scored.
- Competitor scale rewritten rather than lengthened. Scale length was held at three items per style so that scoring stays identical and comparable across all six styles. The three Competitor items now cover three distinct facets of the construct - affect on winning, self-mastery against one's own previous best, and explicit motivation to win - with no novelty language (which pulled one item toward Explorer) and no self-tracking language (which pulled another toward Organizer). Verbatim replacement wording is withheld from the public edition.
- Attention check moved to Tier 1. The instructed-response item was relocated from the premium-only Tier 2 block into the shared Tier 1 Likert block, so data quality can now be flagged for free respondents as well as premium ones. Wording, expected response, and unscored status are unchanged.
Scale length remains three items per style. The evaluation's recommendation to lengthen every scale to four or five items was deliberately not adopted at this stage: keeping all six scales at equal length keeps the composite score directly comparable across styles without reweighting. Competitor scores remain provisional in research use until the revised scale is validated in the next wave.
10.9 Instrument version history and dataset flagging
Because item wording changed, the dataset is no longer a single homogeneous sample. Responses collected before and after this revision are not directly poolable for the four rewritten items or for any Competitor-scale analysis. Rather than reconstruct that boundary from memory later, the boundary is stored in the data itself.
Mechanism. Every row in quiz_results carries an instrument_version
column. All responses collected before this revision are 1.0. Every response
collected after it is stamped 1.1 at submission time by the server, from a
single constant in the codebase. The column is included in the raw CSV export,
so any downstream analysis can split on it directly.
How the effective date is set. The version's effective date is not the day the edits were authored. It is derived at read time from the earliest recorded response under that version, which is necessarily the day the change was published. The admin dashboard shows a version as "Pending publish" until its first response arrives, then displays the derived in-effect date. This makes it impossible for the recorded boundary and the real boundary to drift apart.
Where to see it. Admin dashboard → Instrument versions tab: one entry
per version, with its effective date, response count, and the list of changes.
The registry behind that tab lives in src/lib/instrumentVersions.ts and is the
machine-readable twin of this section.
Maintenance rule (standing). Any change to item wording, keying, scoring, scale length, or item placement requires all three of the following in the same change, or it is incomplete:
- bump
CURRENT_INSTRUMENT_VERSIONand add an entry toINSTRUMENT_VERSIONS; - add a numbered subsection here in section 10 describing what changed and why;
- update the public edition of this document via the marker-based generator.
| Version | Effective | Sample | Summary |
|---|---|---|---|
| 1.0 | Launch through Aug 2026 | Wave 1 validation sample (N = 4,127 submissions; cleaned N = 3,730) | Original instrument. Four reverse-keyed Tier 1 items; attention check premium-only. |
| 1.1 | Set on publish (see dashboard) | Wave 2, collecting | Reverse items removed, Competitor scale rebuilt, attention check moved to shared Tier 1 block. |
11. How these findings shaped the report
| Finding | Design decision |
|---|---|
| Derived opposite pairs (9.1) | Opposite section uses a fixed pairing map, config-locked and identical for all users, distinct from the person-specific growth edge based on lowest style |
| Opposites co-lead below chance (9.2, 9.6) | A dedicated "you carry your own opposite" copy variant for the 16% whose secondary is their primary's opposite |
| Tenfold prevalence spread and direction sensitivity (9.2, 9.6) | Live rarity statistic with ordered-pair matching; banding at rare under 1.5%, uncommon 1.5 to 4%, common above 4%; minimum-count guard at 30 observations; mandatory healthy framing |
| Dimensions are not redundant with styles (9.6) | The two-layer structure is retained: styles and dimensions are reported as separate sections rather than collapsed into one profile |
| Rank ties and blended profiles | Deterministic tiebreak cascade, documented in 7.5.2 |
| Dreamer burnout elevation (section 9) | Dreamer "when play goes missing" copy written with extra care for an audience arriving depleted |
| Organizer playfulness deficit (section 9) | Organizer report explicitly reframes ordering, curating, and hosting as play |
| Gender prevalence skew (section 9) | Current percentages never presented as population facts; demographic weighting planned as the sample grows |
| Attention check | Moved into the Tier 1 block in v1.1 so free and premium respondents are both screened; flagged responses retained and scored; every published figure sensitivity-tested with and without them, and reported both ways where it moves materially |
We include this table because it demonstrates something we think matters: the report changed in response to evidence. Several sections exist, and one was written differently, because of what the data showed.
12. Limitations
- Convenience sample. Self-selected from audiences already interested in play; 64% female; English-language; skewing 25 to 44. Not representative of any general population.
- Rank-only export. (Closed, September 2026.) The opposite-pair derivation has been confirmed against raw composite score correlations on the full sample of 20,966 responses; see section 9.6.
- Cross-sectional self-report. All wellbeing relationships are associations, not causal claims. We do not know whether play reduces burnout, burnout reduces play, or a third factor drives both.
- Compositional constraint. Forced ranks inflate negative correlations uniformly, interpreted against the -0.20 baseline throughout.
- Simulation scope. The pre-launch simulation validates scoring machinery only.
- Short Likert scales. Three items per style limits achievable internal consistency independent of item quality. Confirmed empirically in section 10.2: α = .51 to .64 for the five healthy scales, with the projected fix being length rather than content.
- Derived strengths. Character strengths are assigned by style, not measured.
- Competitor is provisional. In the August 2026 evaluation the Version 1.0 Competitor scale did not hold together (α = .26). It has since been rewritten (section 10.8), but Competitor scores and Competitor primary-style assignments should be treated as provisional in research use until the revised scale is validated on wave 2.
- Model fit below conventional standards. CFI = .787 in the holdout confirmatory model on the Version 1.0 form, before the item revisions. The five healthy scales reproduce their correlations reasonably well on residual-based indices, but the six-factor model as specified does not yet meet the .90 benchmark.
- No test-retest evidence. Retake linkage failed for 388 repeat submissions in the evaluation window, so stability over time is entirely unestablished. Retake linkage is being repaired, and stability will be reported in the wave 2 evaluation.
PART V: WHAT IS OUTSTANDING
13. What we have not yet established
Published deliberately. Everything below is outstanding, ordered by how much it matters. Items that appeared here in earlier versions of this document and have since been completed — internal consistency, item-level factor analysis, item-level discriminant validity, and internal convergent evidence across formats — have been removed from this list and are reported in section 10.
Tier 1: needed before any strong accuracy claim
Test-retest reliability. A subsample retaking after two to four weeks, reporting both the correlation of continuous scores and the percentage whose primary style is unchanged. This is the most important number missing from this document. MBTI's roughly fifty percent type-change rate on retest (Pittenger, 1993, 2005; Salter, Evans, & Forney, 1997) is the most cited criticism of the personality industry, and a published figure of our own, whatever it turns out to be, is a decisive differentiator. It is also the empirical basis for the ninety-day retake cooldown in the paid product. Results history is already stored, which makes this study feasible now.
Re-validation of the revised instrument. The August 2026 evaluation (section 10) ran the exploratory and confirmatory factor analyses on the Version 1.0 form and reported them in full, including the findings that went against the model. Four items were rewritten as a result. The same analysis now has to be repeated on wave 2 data, targeting CFI and TLI at or above .90, before any six-factor structural claim can be made. Until that is done, the Competitor scale in particular remains provisional.
Dimension-to-style correlations. Partly answered. Social Orientation and the Tier 1 context sliders have been correlated against style scores on the full sample (section 9.6); all associations are small, so that layer is not redundant with the styles. The six Tier 2 dimensions still need the same treatment once premium volume allows.
Tier 2: internal evidence, in place of external convergence
We have made a deliberate choice not to administer Play Styles alongside a clinical or established playfulness scale (for example Proyer, 2017, or Shen, Chick, & Zinn, 2014). Play Styles is a self-insight and language tool, not a clinical or diagnostic instrument, so the goal is internal coherence, reliability, and practical usefulness rather than convergence with a validated psychological construct. The absence of a convergent measure is a design decision, not a gap awaiting closure, and this document asks to be evaluated on those terms.
That choice has a cost, and we state it rather than hide it. We cannot answer the question "is this simply an existing personality construct in new language?" with data. We answer it with the argument in section 4: the instrument measures entry points into play, which is not what existing instruments measure, and the structure it predicts shows up in real response patterns (section 9.1 and 9.2). A reader who finds that insufficient is entitled to, and the appropriate response is to weigh the instrument by its usefulness rather than by its correspondence to another scale.
Some of that internal evidence is now in hand. Item-level discriminant validity has been established for the Tier 1 bank: every HTMT ratio falls between −.01 and .53, far below the .85 threshold, and each item's behavior against its own scale is reported item by item in section 10.4. Convergence across the Likert, ranking, and forced-choice formats has also been measured (section 10.5), and criterion evidence against the context sliders is reported in section 10.6.
What remains outstanding in place of external convergence:
The same internal evidence for the Tier 2 subtype bank. Discriminant and convergent analysis of the 48 subtype items — each correlating most strongly with its own subtype, subtype items correlating with their parent style's Tier 1 items, and within-style ranking positions correlating with the corresponding Likert sums. The premium bank has not yet been evaluated at volume.
Criterion evidence from behavior rather than from self-report in the same sitting. The section 10.6 criteria are context sliders answered in the same session, which shares method variance with the items. The stronger test is Live Playfully Journal behavior: daily play frequency, energy change, self-reported play type. That is a better test of practical usefulness than agreement with a second self-report scale would be.
Known-groups comparisons by sector, age band, and partner cohort.
Tier 3: needed for organizational and cross-cultural use
Measurement invariance and differential item functioning, across gender first, given the prevalence differences observed in the July sample, then age, culture, and language.
Criterion validity. Whether Play Style or play deprivation predicts wellbeing, engagement, burnout, or life satisfaction over time. Longitudinal data becomes available as repeat takers accumulate.
Demographically weighted prevalence figures, replacing raw convenience-sample percentages. Note that these remain prevalence within our respondents, not norms against a standardized population; the codebook is explicit that there is no external norm group and no percentiles or cut scores.
14. Appropriate use
Play Styles is designed for self-understanding, conversation, and team development. It supports reflection and shared language. It does not support decisions about people.
Specifically: it should not be used to select, promote, assign, or evaluate anyone. It should not be used to explain away conflict or excuse behavior. It should not be presented as a fixed identity.
PART VI: PRIVACY AND DATA PROTECTION
15. Privacy and data protection rationale
Effective date of the current privacy policy: August 12, 2026. This describes our practices and the reasoning behind them.
15.1 Purpose of this document
This document explains, in one place, how our web products handle personal information and why we made each design choice.
15.2 What the products do
- Play Styles Quiz (free). A self-report questionnaire. Returns six play style scores, a results page, an emailed summary, and a PDF.
- Play Styles Premium Assessment ($49). A longer version of the same question set returning subtype scores, dimensions, and a 12-page report.
- Live Playfully Digital Journal ($35). A 90-day account-based reflection journal with daily entries, sprint reviews, and a capstone. This is a separate product from Play Styles, but it runs in the same web interface, so it is covered by the same privacy practices described here.
All three are consumer-facing, self-initiated, and voluntary.
15.3 Data we collect and why
| Category | Examples | Why we need it |
|---|---|---|
| Identity | First name, last name, email | Deliver results, email the report, let people return to their account |
| Assessment data | Question answers, computed scores | The product itself; aggregated research |
| Optional demographics | Age range, gender, country, sector | Research segmentation; every field is optional |
| Journal content | Daily entries, reflections | The product itself; visible only to the account owner |
| Purchase records | Stripe session/payment IDs, product, amount | Grant access, support, accounting |
| Technical | IP address, user agent, timestamps | Security, abuse prevention, fraud checks |
| Attribution | Referral code from a referral link; UTM parameters and referring domain (hostname only) of the first page visited; landing path | Credit referrals; understand which channels bring people to the assessment. First-party only; no full referrer URL, no search terms, no third-party pixels |
| Assessment progress | Pseudonymous session id, furthest step reached, timestamps, attribution fields above | Understand where people stop before finishing. Held in a separate table from results and never merged into the analysis dataset |
| Referral participation | Referrer's name, email, payout details, and the count of qualifying purchases made through their link | Operate the referral program and pay the flat $8 per-purchase commission. Referrers never see who used their link |
| Reviews and testimonials | Rating, written review, display name if given | Product improvement; published only with the reviewer's consent |
| Contact preferences | Email address, marketing opt-in state, unsubscribe timestamp | Deliver requested email and, after an unsubscribe, keep the address on a suppression list so we can honor the opt-out |
We do not collect: card numbers (Stripe handles payment entirely), precise location, contacts, biometrics, or government identifiers.
Third-party links. The physical Play Journal is sold through an external shop. Following that link leaves our site, and anything entered there is governed by that shop's own privacy policy; we receive no order data from it.
Research use. Research and publication outputs use aggregated, de-identified data only. Individual answers are never published or shared as an individual record.
15.4 Legal bases (GDPR / UK GDPR)
- Contract — delivering the results, report, and journal a person requested.
- Consent — marketing email, research use of individual responses, and optional referral attribution storage.
- Legitimate interests — security, fraud prevention, and abuse mitigation, balanced against a low-risk, self-submitted dataset.
For US state privacy laws (CCPA/CPRA, VCDPA, CPA and similar), the operative facts are: we do not sell personal information, we do not share it for cross-context behavioral advertising, and we honor access and deletion requests through a self-serve control in the product.
15.5 Sensitivity assessment (why this is not health data)
Some items touch on energy, stress, and well-being. We deliberately do not treat, market, or use these results as health, clinical, diagnostic, insurance, or employment information, and the privacy policy states this explicitly. The outputs are reflective and research-oriented. This keeps the dataset outside "health data" definitions in laws such as Washington's My Health My Data Act and the special-category rules in GDPR Art. 9, and it is enforced in practice by (a) the wording of the questions, (b) the disclaimer on the policy page, and (c) the fact that no third party receives individual-level results.
15.6 Minimum age: 16
We set a single global floor of 16 years old for taking an assessment, purchasing, or creating an account. Reasoning:
- COPPA (US) applies under 13.
- GDPR Art. 8 lets member states set the digital-consent age anywhere from 13 to 16; several set it at 16.
- Rather than build per-country age logic and a parental-consent workflow, we adopted the strictest common threshold. At 16+, no parental consent mechanism is required anywhere we operate.
Implementation: the age dropdown's lowest option is "16–17" (the previous "Under 18" option was removed), and the rule is stated in both the Privacy Policy and the Terms.
15.7 Cookies and browser storage
We use only first-party cookies and browser storage. There are no advertising pixels, no ad networks, no cross-site trackers, and no data brokers.
| Item | Category | Purpose |
|---|---|---|
| Auth session | Essential | Keeps a signed-in user signed in |
| Quiz progress | Essential | Saves an in-progress assessment on the device (~30 days) |
| Session id for drop-off | Essential | Pseudonymous id so an abandoned session is counted once, not repeatedly |
| Stripe fraud protection | Essential | Runs only on checkout pages |
| Referral code | Essential | Attributes a purchase to the referral link the visitor arrived through, and applies that link's 10% discount at checkout (~60 days). Strictly necessary to deliver the discounted price the visitor was offered |
| Traffic source (UTM / referring domain) | Essential | First-touch record of which link or channel brought the visitor in (~60 days). Hostname only, first-party, never shared |
| Consent choice | Essential | Remembers the banner answer so we stop asking |
A public, plain-language /cookies page lists each item, its category, and its
retention. A non-blocking consent banner offers "Essential only" or "Allow
referral tracking"; declining prevents the optional items from ever being
written. Essential storage is not gated because it is strictly necessary to
deliver a service the user requested — the standard exemption under the ePrivacy
Directive and equivalent guidance.
15.8 Do Not Track / Global Privacy Control
Because we run no cross-site tracking and sell no data, a DNT or GPC signal does not change our behavior — there is nothing to switch off. CalOPPA requires that we disclose how we respond to such signals; the Privacy Policy contains that statement rather than claiming a capability we don't need.
15.9 Retention and de-identification
- Free and premium assessment records are automatically de-identified after five years: first name, last name, and email are stripped, anonymous scores remain for research.
- Account and journal data persist while the account is active.
- On account deletion, personal details are removed promptly and only de-identified scores are retained, for both free and premium assessments.
- Anonymous, aggregated research data is retained indefinitely and is no longer personal information.
De-identified scores surviving indefinitely is what makes long-run research possible. The research value and the privacy commitment come from the same design decision: the record stays, the person's identity does not.
15.10 Individual rights and how we honor them
Signed-in users have two self-serve controls in their account page:
- Download my data — a machine-readable JSON export of every record tied to the account: quiz results, premium results, journal profile, onboarding, daily entries, sprint reviews, capstone, and purchases.
- Delete my account — permanently deletes the login and all journal content, and de-identifies both free and premium assessment rows — name, email, and the link to the account are removed — so anonymous research scores survive with no link to the person. Treatment of the two assessment tiers is identical.
Both run server-side and are scoped to the verified session user, so one person can never reach another's data. Requests can also be made by email to inquiry@nifplay.org, including access, correction, deletion, and unsubscribe. Unsubscribe links appear in every marketing email.
15.11 Security controls
- Data is stored in a managed, encrypted Postgres database (Supabase) with encryption in transit and at rest.
- Row Level Security is enabled on every user-facing table, so a signed-in user can only read and write their own rows.
- Quiz submissions are written exclusively through a validated server function using a privileged key; the public/anonymous role has no insert path.
- All submitted payloads are schema-validated and size-capped before they touch the database, which blocks malformed and oversized input.
- Privileged database functions have execute permission revoked from the anonymous and authenticated roles.
- Administrative views require an explicit admin role stored in a dedicated roles table, checked server-side — never in the browser and never in a user-editable field.
- Payments never touch our servers; Stripe is PCI-DSS compliant and we receive only non-sensitive confirmation identifiers.
- Automated security scanning runs against the backend, and findings are tracked and remediated.
15.12 Processors and sub-processors
| Vendor | Role | Data seen |
|---|---|---|
| Supabase | Database, authentication, storage | All application data |
| Stripe | Payment processing | Name, email, payment details |
| Email delivery provider | Transactional and results email | Name, email, results summary. Sends are queued durably and retried, so a delivery failure does not lose the message |
| Hosting/CDN | Serving the site | IP address, request metadata |
Each acts as a processor on our behalf under its standard data processing terms. None is authorized to use the data for its own purposes.
15.13 International transfers
Participants come from many countries; our infrastructure is US-based. Transfers rely on the processors' Standard Contractual Clauses and equivalent transfer mechanisms in their data processing agreements.
15.14 Research use
Individual-level responses are used to deliver a person's own results. Research outputs use aggregated, de-identified data only. No individual record is published, sold, or shared with a partner organization. Consent for research use is obtained at submission and can be withdrawn by deleting the account.
16. References
First pass, compiled August 2026. Each entry should be verified against the source and matched to an in-text citation before publication.
Affective neuroscience and the primary emotional systems
- Panksepp, J. (1998). Affective Neuroscience: The Foundations of Human and Animal Emotions. Oxford University Press.
- Panksepp, J., & Biven, L. (2012). The Archaeology of Mind: Neuroevolutionary Origins of Human Emotions. W. W. Norton.
- Davis, K. L., & Panksepp, J. (2018). The Emotional Foundations of Personality: A Neurobiological and Evolutionary Approach. W. W. Norton.
- Davis, K. L., Panksepp, J., & Normansell, L. (2003). The Affective Neuroscience Personality Scales: Normative data and implications. Neuropsychoanalysis, 5(1), 57–69.
Play in the brain and across species
- Vanderschuren, L. J. M. J., Achterberg, E. J. M., & Trezza, V. (2016). The neurobiology of social play and its rewarding value in rats. Neuroscience & Biobehavioral Reviews, 70, 86–105. https://doi.org/10.1016/j.neubiorev.2016.07.025
- Achterberg, E. J. M., & Vanderschuren, L. J. M. J. (2023). The neurobiology of social play behaviour: Past, present and future. Neuroscience & Biobehavioral Reviews, 152, 105319.
- Pellis, S. M., & Pellis, V. C. (2009). The Playful Brain: Venturing to the Limits of Neuroscience. Oneworld.
- Burghardt, G. M. (2005). The Genesis of Animal Play: Testing the Limits. MIT Press.
- Panksepp, J. (2007). Can PLAY diminish ADHD and facilitate the construction of the social brain? Journal of the Canadian Academy of Child and Adolescent Psychiatry, 16(2), 57–66.
Play in human development and adult life
- Brown, S., with Vaughan, C. (2009). Play: How It Shapes the Brain, Opens the Imagination, and Invigorates the Soul. Avery.
- Gray, P. (2011). The decline of play and the rise of psychopathology in children and adolescents. American Journal of Play, 3(4), 443–463.
Measurement of adult playfulness
- Proyer, R. T. (2017). A new structural model for the study of adult playfulness: Assessment and exploration of an understudied individual differences variable. Personality and Individual Differences, 108, 113–122. https://doi.org/10.1016/j.paid.2016.12.011
- Proyer, R. T., & Jehle, N. (2013). The basic components of adult playfulness and their relation with personality: The hierarchical factor structure of seventeen instruments. Personality and Individual Differences, 55(7), 811–816.
- Shen, X. S., Chick, G., & Zinn, H. (2014). Validating the Adult Playfulness Trait Scale (APTS). American Journal of Play, 6(3), 345–369.
Motivation and wellbeing
- Deci, E. L., & Ryan, R. M. (1985). Intrinsic Motivation and Self-Determination in Human Behavior. Plenum Press.
- Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68–78.
Personality frameworks and their critics
- Jung, C. G. (1971). Psychological Types (Collected Works, Vol. 6; original work published 1921). Princeton University Press.
- Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) and NEO Five-Factor Inventory: Professional Manual. Psychological Assessment Resources.
- Pittenger, D. J. (1993). The utility of the Myers-Briggs Type Indicator. Review of Educational Research, 63(4), 467–488.
- Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221. https://doi.org/10.1037/1065-9293.57.3.210
- Salter, D. W., Evans, N. J., & Forney, D. S. (1997). Test-retest of the Myers-Briggs Type Indicator: An examination of dominant functioning. Educational and Psychological Measurement, 57(4), 590–597.
Character strengths
- Peterson, C., & Seligman, M. E. P. (2004). Character Strengths and Virtues: A Handbook and Classification. Oxford University Press and American Psychological Association.
Measurement method and psychometrics
- Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22(140), 1–55.
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334.
- McDonald, R. P. (1999). Test Theory: A Unified Treatment. Lawrence Erlbaum Associates.
- Jarvis, C. B., MacKenzie, S. B., & Podsakoff, P. M. (2003). A critical review of construct indicators and measurement model misspecification in marketing and consumer research. Journal of Consumer Research, 30(2), 199–218.
- Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research. Journal of Applied Psychology, 88(5), 879–903.
- Oppenheimer, D. M., Meyvis, T., & Davidenko, N. (2009). Instructional manipulation checks: Detecting satisficing to increase statistical power. Journal of Experimental Social Psychology, 45(4), 867–872.
About this document
Author: Mia Sundstrom, National Institute for Play Version: 1.0 Published: August 2026
This document is versioned and will be updated as evidence accumulates. Section 13 lists what remains outstanding, and each item that is completed will be reported here whatever the result.
Questions, corrections, and requests to review the underlying data can be sent to inquiry@nifplay.org.
Suggested citation: Sundstrom, M. (2026). Play Styles: Grounding, Design, and Validity (Version 1.0). National Institute for Play.
