Proposed Protocol for IIV Computation in Adolescent Cognitive Monitoring: The PRECISE-ICS Standard
The science of intra-individual variability is validated. The clinical protocol for adolescents doesn't exist yet. This paper proposes one.
Proposed Protocol for Intra-Individual Variability (IIV) Computation in Adolescent Longitudinal Cognitive Monitoring: The PRECISE-ICS Standard
Abstract
Intra-individual variability (IIV) in cognitive performance has accumulated decades of empirical support as one of the most sensitive early markers of neural change available to cognitive science. Longitudinal studies in aging, traumatic brain injury, and neurodegeneration have established that moment-to-moment and session-to-session fluctuation in reaction time and executive function predicts cognitive decline trajectories with greater fidelity than cross-sectional mean performance, often detecting meaningful change years before clinical symptoms emerge. Despite this evidentiary base, a 2025 review in The Clinical Neuropsychologist states plainly that guidelines for calculating, interpreting, and implementing IIV measures in clinical practice are currently lacking. That sentence identifies a real and consequential gap. The science is well-supported. The operational protocol, at least for adolescent ecological contexts, does not yet exist.
This paper proposes the PRECISE Intra-individual Cognitive Stability Standard (PRECISE-ICS): a formal, reproducible methodology for computing, interpreting, and longitudinally monitoring IIV in adolescent populations using browser-based cognitive assessment. The protocol addresses four methodological dimensions that the existing literature leaves unresolved: how IIV should be calculated across reaction time and executive function tasks with explicit justification for method selection from among the five primary approaches in the comparative literature; how raw response data should be cleaned to distinguish genuine signal from browser-based artifact; how a stable personal baseline should be formed longitudinally, including minimum trial and session requirements and a rolling window specification; and how meaningful cognitive drift should be defined and flagged relative to an individual’s own historical stability profile rather than population norms.
The protocol additionally addresses the specific developmental characteristics of the adolescent brain that make standard adult IIV frameworks inapplicable, and proposes a structured validation research agenda for prospective empirical testing of every proposed threshold.
All proposed thresholds in this protocol are operationally derived rather than empirically validated at the time of writing. This distinction is stated at the outset and repeated throughout. The value of this paper is not in producing a finished instrument. It is in producing a rigorous, transparent, and testable starting point for a field that currently has no shared methodological foundation for ecological IIV monitoring in youth populations.
Keywords: intra-individual variability, cognitive monitoring, adolescent cognition, ecological assessment, longitudinal tracking, reaction time variability, IIV protocol, browser-based neuroscience, PRECISE, prefrontal maturation, intra-individual standard deviation, coefficient of variation
Table of Contents:
- Introduction
- Background and Scientific Foundation
- The PRECISE-ICS Protocol Specification
- Adolescent-Specific Protocol Adaptations
- Statistical and Psychometric Considerations
- Browser-Based Timing: A Full Treatment of Limitations and Mitigations
- Proposed Validation Research Agenda
- Ethics and Data Governance
- Conclusion
1. Introduction
1.1 The Signal That Mean Performance Conceals
There is a category of information about cognitive function that aggregate performance scores systematically discard. When a participant completes thirty reaction time trials and produces a mean of 320 milliseconds, that number describes the center of a distribution. It says nothing about the shape of that distribution. Nothing about whether the trials clustered tightly around the mean or scattered unpredictably across a 200-millisecond range. Nothing about whether performance was stable across the session or drifted in a consistent direction. Nothing about whether any individual trials were catastrophically slow in a way that the mean absorbed without registering. The mean is a compression. And like all compressions, it loses information.
The information it loses is not incidental. It is, according to a substantial and growing body of cognitive neuroscience, among the most diagnostically valuable information that reaction time data contains.
This is the foundational argument for studying intra-individual variability. IIV refers to the trial-to-trial and session-to-session fluctuation in an individual’s cognitive performance, considered not as noise to be averaged away but as a signal to be measured and interpreted in its own right. The theoretical account for why this signal carries predictive weight is coherent and well-grounded. A neural system that is functioning normally and whose attentional regulation mechanisms are intact produces consistent outputs. When the systems underlying processing speed, attention, and executive control are operating as expected, reaction times cluster tightly, lapses are infrequent, and session-to-session variability is low and predictable.
When those systems are disrupted, by injury, disease progression, fatigue accumulation, developmental instability, or any other source of neural dysregulation, the outputs become inconsistent.
Variability increases. The distribution widens. Lapses become more frequent. The brain, in this framework, announces its own instability through the shape of its output distribution before it announces it through any change in average speed.
The empirical case for this argument is substantial, though most of it comes from adult and older adult populations, a limitation that Section 2.3 addresses directly. Hultsch, Strauss, Hunter, and MacDonald (2002) followed a sample of older adults across a six-year longitudinal period, measuring both mean reaction time and intraindividual variability repeatedly across that interval. Their finding has become one of the most cited in the IIV literature: individuals with higher baseline IIV at enrollment showed accelerated cognitive decline over the subsequent six years, even when their mean performance at baseline was comparable to individuals who declined slowly. The variability was the leading indicator. The mean was not. MacDonald, Hultsch, and Dixon (2003), extending this line of work within the Victoria Longitudinal Study, demonstrated that within-person fluctuations in cognitive performance carry more diagnostic information than cross-sectional comparisons against population averages, a finding that has direct implications for how monitoring systems should be designed. Duchek and colleagues (2009) showed that IIV measures can distinguish individuals with very mild Alzheimer’s disease from healthy older adults even when mean performance scores appear similar between groups, a result that underscores how much information mean-based assessment discards. In the context of mild traumatic brain injury, Mullen and colleagues (2014) documented that clinically measured simple reaction time was among the most sensitive early markers of post-concussive cognitive disruption, detectable in the acute post-injury window when standard neurological evaluations and self-reported symptom scores remained near-normal.
These findings, taken together, support a proposition that this paper treats as its scientific foundation: IIV is not a methodological nuisance. It is a primary signal. And any cognitive monitoring system that reports only mean performance while discarding variability information is leaving the most sensitive part of the data on the table.
1.2 The Problem This Paper Addresses
If the scientific case for IIV is this strong, the natural question is why it has not been translated into standard clinical and ecological practice. The answer is not theoretical. It is methodological.
A 2025 review in The Clinical Neuropsychologist states the problem directly: guidelines for calculating, interpreting, and implementing IIV measures in clinical practice are currently lacking. A companion observation from the same period notes that IIV is an important but underappreciated aspect of cognitive assessment, and that it may be overlooked in part because no single best method for computing it has been established. These sentences describe not a gap at the frontier of research but a gap between validated science and usable implementation. The science of IIV is not nascent. The field has five established computation methods, each with a legitimate theoretical basis and a set of tradeoffs that make it more or less appropriate for different contexts. What the field lacks is a principled framework for choosing between those methods in any given deployment context, cleaning the data appropriately before applying them, specifying how many trials and sessions are required for stable estimates, and determining when a longitudinal change in IIV is meaningful rather than noisy.
This is the gap. It is not a small one. Its consequences are practical.
Without a shared methodological framework, IIV findings across studies are difficult to compare directly. Researchers adopting IIV-adjacent measurement tools have no principled reference standard for their methodological choices. Consumer-facing and ecological tools that attempt to implement IIV monitoring, and there are now several that do, have no defensible protocol to align to. And in the specific context that this paper addresses, which is browser-based longitudinal monitoring of healthy adolescents, the gap is compounded by two additional problems: the absence of adolescent-specific IIV norms in the published literature, and the structural timing limitations of browser-based assessment environments that introduce artifact classes that laboratory protocols do not encounter.
Existing tools with IIV sensitivity address different parts of this gap but not all of it. Cogstate has strong clinical trial validation in adult populations and uses adaptive paradigms that are sensitive to within-session performance changes, but it is not designed for ecological deployment in non-clinical adolescent settings and its normative base is adult-centered. CNS Vital Signs has been used in school-based concussion monitoring programs but similarly relies on adult-normed population comparisons. The ImPACT battery, the most widely used concussion assessment instrument in athletic settings, has established test-retest reliability in high school athletes across annual intervals (Schatz et al., 2006), but it is an episodic baseline-and-retest system rather than a continuous longitudinal monitoring framework. None of these instruments addresses the specific problem of computing, interpreting, and acting on session-level IIV estimates in a healthy adolescent population monitored repeatedly over months through a browser-based interface.
1.3 The Scope and Contribution of This Paper
This paper proposes the PRECISE Intra-individual Cognitive Stability Standard as a methodological contribution toward closing this gap. PRECISE-ICS is not a finished clinical protocol. It does not have prospective validation data behind it at the time of writing. What it has is a rigorous, explicit, and testable set of methodological choices that are derived from the comparative IIV literature, the cognitive neuroscience of adolescent brain development, the psychometric requirements for stable longitudinal estimation, and the specific engineering constraints of browser-based assessment environments.
The contribution is to make those choices public, explicit, and arguable. A methodology proposal that hides its assumptions behind vague language about best practices cannot be tested, critiqued, or improved. A methodology proposal that names specific thresholds, states the reasoning behind each one, and acknowledges exactly where the reasoning runs ahead of the available evidence is a contribution that invites the kind of scientific engagement that leads to eventual validation and refinement.
Every threshold in this document that is operationally proposed rather than empirically validated is marked as such. The distinction matters and it is never elided.
2. Background and Scientific Foundation
2.1 What IIV Actually Measures: A Theoretical Account
Before specifying how to compute IIV, it is worth being precise about what it reflects biologically, because the theoretical account determines which computation method is appropriate for which question.
The cognitive neuroscience of IIV points to two partially separable sources of trial-to-trial variability in reaction time data.
The first is moment-to-moment fluctuation in attentional engagement, sometimes called the tonic-phasic attention framework. Attention is not a sustained constant state. It fluctuates. Periods of high engagement produce faster, more consistent responses. Attentional lapses, which are brief episodes in which the attentional system disengages from the task, produce isolated slow responses that inflate variability without necessarily reflecting a change in the person’s average processing capacity. This is the source of variability that MSSD is particularly sensitive to, and it is what ex-Gaussian tau is theoretically designed to capture through the shape of the slow tail of the RT distribution.
The second source is neural noise, the baseline level of stochastic variability in signal transmission across the networks underlying the task. Neural noise is higher when neural integrity is compromised, whether by age-related synaptic degeneration, acute injury disrupting axonal conduction, or developmental immaturity of myelination and synaptic organization. This is the source of variability that the IIV literature’s long-term predictive findings are primarily detecting: the fact that individuals with higher baseline IIV go on to show accelerated decline reflects the fact that higher neural noise at baseline is a marker of reduced neural reserve, not just a behavioral curiosity.
These two sources interact in practice and cannot be fully separated with behavioral data alone. But understanding that they are distinct influences the interpretation of IIV signals. A session with high variability driven primarily by attentional lapses, which show up as isolated extreme-slow trials against a background of otherwise normal performance, looks different in the data from a session with high variability driven by elevated neural noise, which produces a uniformly wider distribution without extreme outliers. The computation method chosen determines which of these patterns the IIV estimate is most sensitive to, and this is why method choice is a theoretical as well as a practical decision.
2.2 The Five Primary Computation Methods: A Comparative Analysis
The literature has converged on five primary methods for quantifying IIV in reaction time tasks. The comparative analysis presented here draws on published methodological reviews and constitutes the scientific basis for the method selections in Section 3.
Raw intra-individual standard deviation (iSD) is computed as the standard deviation of reaction times across trials within a session for a given individual. It is the simplest and most intuitive estimate of within-person variability and has the longest history in the IIV literature. Its central limitation is the absence of any correction for mean speed differences. Because the absolute spread of a distribution tends to scale with its mean in RT data, individuals who are systematically slower will tend to produce larger raw iSD values even if their proportional variability is identical to faster individuals. This creates a confound between mean performance level and variability that makes raw iSD difficult to interpret in contexts where mean RT is expected to shift, which includes longitudinal monitoring across development, across recovery, and across sessions that differ in testing conditions.
Residualized iSD addresses this limitation by ‘partialling’ out systematic between-subject and within-subject variance through linear mixed-effects modeling before computing the standard deviation of the residuals. The residualized estimate reflects only the variability that remains after accounting for mean-level differences both between individuals and across sessions within individuals. This is the most statistically rigorous approach in the literature and produces the most interpretively clean IIV estimate. It is also computationally impractical for real-time or near-real-time browser-based monitoring. Fitting an appropriate multilevel model requires access to the full historical dataset for each participant at the time of each new session, a computational and architectural requirement that is incompatible with a browser-based system operating without server-side data storage. This is a hard constraint, not a preference, and it eliminates residualized iSD from consideration as a primary metric for the PRECISE-ICS protocol.
Coefficient of Variation (CoV) is computed as the ratio of iSD to the intraindividual mean reaction time for the same set of valid trials. It is a proportional normalization: CoV expresses variability as a fraction of the individual’s own speed, making it interpretable across individuals and sessions regardless of mean RT differences. A person with a mean RT of 400 milliseconds and an iSD of 80 milliseconds has the same CoV of 0.20 as a person with a mean RT of 250 milliseconds and an iSD of 50 milliseconds. This normalization corrects for the most consequential confound in within-person longitudinal monitoring: mean speed drift over time, across sessions, and across development. CoV does not correct for within-session mean trends, which is a limitation addressed in the protocol specification. It requires approximately 30 or more valid trials per session to produce a stable estimate, a threshold derived from the psychometric literature on RT variability reliability. Below this threshold, a single outlier trial can swing the iSD substantially, and the resulting CoV estimate becomes too noisy to serve as a reliable longitudinal reference point.
Mean Square of Successive Differences (MSSD) is computed as the mean of the squared differences between each adjacent trial pair in sequence:
MSSD = (1 / (n-1)) * sum[(RT_{i+1} - RT_i)^2]
Unlike iSD and CoV, which are computed across all trials simultaneously and treat the trial sequence as unordered, MSSD is inherently sequential. It captures moment-to-moment fluctuation by measuring how much each trial differs from the one immediately preceding it. MSSD is therefore particularly sensitive to attentional lapses, because a lapse produces a large trial-to-trial difference both on the lapse trial itself (which is slow relative to the preceding trial) and on the recovery trial (which is fast relative to the lapse). A distribution with the same mean and standard deviation as another but with clustered slow trials will produce a higher MSSD than one whose slow trials are evenly dispersed, capturing something real about the attentional dynamics that the standard deviation misses.
The relationship between MSSD and iSD depends critically on the autocorrelation structure of the data. When autocorrelation in the RT sequence approaches zero, meaning that each trial’s speed is independent of the preceding trial’s speed, MSSD and iSD converge to approximately the same quantity. When positive autocorrelation is present, meaning that slow trials tend to cluster together, MSSD and iSD diverge, with MSSD becoming relatively larger and more informative about lapse clustering specifically. This dependence on autocorrelation structure means that MSSD’s interpretive value varies across task designs and participant states, making it less suitable as a primary metric but highly valuable as a supplementary lapse-detection index when trial counts are sufficient.
MSSD is also scale-dependent in the same way raw iSD is: individuals with higher mean RT will tend to produce higher raw MSSD values. This can be partially addressed by computing a normalized MSSD as the ratio of MSSD to the squared mean RT, analogous to the normalization CoV applies to iSD. The protocol specifies this normalization for MSSD when it is used as a supplementary metric.
Ex-Gaussian tau is the exponential component of the ex-Gaussian distribution, which decomposes the RT distribution into a normal component (described by mu and sigma) and an exponential component (described by tau). The exponential component captures the positive skew of RT distributions, which reflects the slow tail of responses associated with attentional lapses and processing failures. Tau has strong theoretical grounding as a specifically attentional index: increases in tau are associated with increased lapse frequency without necessarily reflecting changes in mean processing speed or distribution width. It is the most theoretically precise of the five methods for isolating the attentional lapse component of IIV.
It is also the most data-hungry. Published recommendations for stable ex-Gaussian parameter estimation generally require a minimum of 80 valid trials per session, with some analyses suggesting 100 or more for reliable tau estimates. This is a prohibitive constraint for ecological deployment where session brevity is essential for repeated engagement. A protocol that requires 100 trials per session to compute its primary IIV metric is not an ecological protocol. Ex-Gaussian tau is therefore excluded from the PRECISE-ICS protocol as a primary or supplementary metric for current deployment, with the explicit acknowledgment that it should be revisited if future PRECISE task designs can accommodate the required trial counts within acceptable session durations.

Image Credit: ChatGPT
2.3 Why Adolescent IIV Requires Its Own Framework
The existing IIV literature is overwhelmingly adult-centered.
Its normative foundations, its predictive validity studies, its longitudinal findings, all of them are built on adult or older adult samples. The application of adult IIV frameworks to adolescent populations is not merely an extrapolation. This paper argues it is a category error, because the adolescent brain is doing something that the adult brain is not: it is actively reorganizing itself in ways that are documented to change the baseline level and trajectory of IIV.
The structural basis for this developmental IIV trajectory is well-documented in the literature. Synaptic pruning in the prefrontal cortex continues throughout the second decade of life. The elimination of excess synaptic connections in this region progressively enhances the signal-to-noise ratio in neural processing, supporting the complex integration required for cognitive control (Luna et al., 2004). Myelination of long-range cortico-cortical connections, which increases the speed and reliability of neural signal transmission, also continues into the mid-twenties. Tamnes and colleagues (2010, 2017) documented this structural maturation across convergent longitudinal samples, confirming that cortical thinning in the lateral prefrontal cortex, a region central to executive control and attentional regulation, follows a protracted developmental timeline extending well into the twenties.
The behavioral consequence of this structural trajectory is directly relevant to IIV measurement. As Luna and colleagues established, the refinement of higher-order cognition in adolescence is evidenced specifically by a reduction in the variability of performance on cognitive tasks, including reduced error rates and the stabilization of response latency. The transition from adolescent to adult cognition is, in their framework, a process of increasing the reliability of successful cognitive control and reducing behavioral variability, leading to consistently optimal responses. Inhibitory control in particular, the ability to suppress a prepotent response in favor of a planned one, shows a clear developmental trajectory through adolescence, with error rates decreasing and response consistency increasing as prefrontal networks mature.
What this means for an IIV monitoring protocol is that adolescent IIV baselines are not stable reference points.
They are moving targets, trending downward as a function of normal development.
An IIV value that would be interpretively neutral in a 25-year-old may be entirely developmentally expected in a 14-year-old. A declining IIV trend across sessions in an adolescent user may reflect nothing more than normal prefrontal maturation, which is progress, not pathology. And a stable or rising IIV trend against a background in which developmental reduction is expected may represent a meaningful signal: the brain is not doing what a maturing brain should be doing, which is becoming more consistent.
A protocol that applies static adult normative thresholds to an adolescent monitoring context is, this paper argues, applying the wrong reference standard for the reasons documented above. The PRECISE-ICS protocol addresses this by rejecting population-normed thresholds as the primary interpretive framework and replacing them with individual-referenced longitudinal modeling, while additionally specifying developmental accommodations for how the rolling baseline and drift detection system should respond to the expected directional trend of IIV change in this age group.

Image Credit: ChatGPT
2.4 The Ecological Validity Problem
Laboratory cognitive assessment achieves its precision partly through a set of controls that ecological deployment cannot replicate: standardized hardware with known timing characteristics, controlled lighting and noise environments, trained administrators who monitor for off-task behavior, and participant populations that have been screened for the conditions the assessment is designed to study. Browser-based ecological assessment gives up most of these controls in exchange for the deployability and frequency of measurement that make longitudinal IIV monitoring possible.
This tradeoff is not a bug. It is a deliberate design choice. The entire premise of ecological IIV monitoring is that frequent measurement in natural environments produces longitudinal data that no laboratory protocol can generate, because no participant will come into a laboratory every week for a year. The longitudinal signal that emerges from ecological deployment is inaccessible through laboratory methods. But the tradeoff has methodological consequences that a rigorous protocol must address explicitly rather than ignore.
The most significant consequence for IIV measurement specifically is that browser-based environments introduce sources of variable latency, meaning trial-to-trial differences in timing precision, that laboratory systems control for but browser systems cannot fully eliminate. Variable latency inflates iSD and therefore CoV without reflecting any change in the participant’s cognitive state. A protocol that does not address this source of artifact will produce IIV estimates that confound genuine cognitive variability with measurement noise.
The PRECISE-ICS Standard addresses this through the data cleaning rules specified in Section 3.3 and the timing architecture specifications in Section 6. The combination of these two protocol components does not eliminate timing noise, but this paper proposes that it reduces it to a level that is small relative to the IIV thresholds used to flag meaningful drift. Whether that claim holds empirically is one of the questions the validation agenda is designed to test.
3. The PRECISE-ICS Protocol Specification
3.1 Architectural Overview
The PRECISE-ICS Standard is structured around four operational components, each of which addresses a distinct methodological problem. The calculation methods component specifies which IIV metrics are computed and why, including the justification for each choice in the comparative literature. The data cleaning component specifies how raw response data is processed before metric computation, including rules for excluding artifact classes that are specific to browser-based environments. The baseline formation component specifies the minimum requirements for establishing a stable personal IIV reference, including trial counts, session counts, and the rolling window structure. The drift detection component specifies how the system determines whether a new session’s IIV represents a meaningful departure from the individual’s personal baseline.
These four components are designed as an integrated system. The calculation methods are only meaningful given the data cleaning rules that define what counts as a valid trial. The baseline formation requirements are only meaningful given the calculation methods that determine how much data is needed for a stable estimate. The drift detection thresholds are only meaningful given the baseline formation rules that define what the reference distribution represents. No component should be adopted in isolation from the others.
3.2 IIV Calculation Methods
3.2.1 Primary Metric for Reaction Time Tasks: Coefficient of Variation
The primary IIV metric for all PRECISE reaction time tasks is the Coefficient of Variation, computed as:
CoV = iSD / M_RT
where iSD is the intraindividual standard deviation of valid reaction times within a session after data cleaning, and M_RT is the intraindividual mean reaction time for the same set of valid trials.
The selection of CoV over the other available methods requires explicit justification, because the choice is consequential and should not be treated as obvious.
CoV is selected over raw iSD because mean RT is not stable across adolescent development, across sessions at different times of day, across different fatigue states, or across the wide range of device types on which browser-based assessment runs. A monitoring system that uses raw iSD as its primary variability metric will systematically confound changes in processing speed with changes in consistency. If an adolescent participant’s mean RT decreases by 30 milliseconds over six months of normal development, and their raw iSD remains constant, the raw iSD will appear unchanged while CoV will correctly register an increase in proportional variability. Conversely, if both mean RT and iSD decrease proportionally, raw iSD will flag this as improvement in variability when CoV correctly identifies it as a stable consistency profile at a faster speed. The normalization CoV applies is not cosmetic. It separates the two questions that cognitive monitoring most needs to keep separate: how fast is this person, and how consistent are they.
CoV is selected over residualized iSD because residualized estimation requires access to the full historical dataset for each participant at the time of each new session, which demands server-side data infrastructure that conflicts with the PRECISE system’s local-only data processing architecture. This is a hard constraint with an ethical dimension: storing raw participant response data on a server introduces privacy exposure that the PRECISE system is designed to avoid. The statistical rigor that residualized iSD offers is real, but it cannot be purchased at the cost of the privacy architecture that makes repeated longitudinal participation ethically defensible.
CoV is selected as the primary metric over MSSD because, as the comparative analysis in Section 2.2 establishes, MSSD’s interpretive value depends critically on the autocorrelation structure of the trial sequence, which cannot be assumed to be consistent across sessions or participants. CoV’s behavior is more predictable across variable data structures, making it a more reliable primary metric for a monitoring system that must function consistently across a heterogeneous population of users.
The limitation of CoV that must be stated clearly and returned to throughout interpretation is that it does not correct for within-session mean trends. If a participant’s RT systematically increases across the 30 trials of a session due to within-session fatigue, the iSD will be inflated by the linear trend component rather than reflecting only the random trial-to-trial fluctuation around a stable mean. This is the reason the protocol tracks the within-session fatigue index, computed as the percentage change in mean RT from the first eight trials to the final eight trials, as a separate quality-of-estimate indicator. The fatigue index operates at two proposed thresholds that serve different purposes. When it exceeds 20 percent, the session’s CoV estimate is flagged as potentially trend-inflated and the associated uncertainty is noted in the longitudinal record, but the session may still contribute to baseline updating if all other validity criteria are met. When it exceeds 35 percent, the trend inflation is considered too severe for the CoV estimate to be interpretable as a stable variability measure, and the session is classified as low-confidence and excluded from baseline updating regardless of other criteria. Prospective validation should examine whether these thresholds are appropriately placed and whether within-session linear detrending prior to CoV computation would reduce the need for fatigue-based session exclusions.
The minimum valid trial count for a reportable session-level CoV estimate is proposed at 30 trials after data cleaning. This threshold is not arbitrary, but it is also not derived from a reliability study conducted in this specific population and deployment context. The psychometric literature on RT variability reliability in laboratory settings suggests that the session-level iSD estimate becomes meaningfully unstable below approximately 30 trials, because a single outlier trial can produce a 10 to 15 percent swing in the standard deviation that would not persist with a larger sample. Whether this threshold transfers accurately to browser-based adolescent assessment is an empirical question that the validation agenda must address. Until that evidence is available, 30 trials is the proposed operational floor, and CoV estimates based on fewer valid trials are reported as low-confidence and excluded from baseline updating.
3.2.2 Supplementary Metric: Normalized MSSD for Lapse Detection
MSSD is designated a supplementary metric in this protocol rather than a primary one, for the reasons established in the comparative analysis. It is retained because it captures something that CoV does not: the sequential clustering of attentional lapses. A session in which slow trials appear in isolated clusters, indicating episodic attentional disengagement, produces a different MSSD profile than a session in which slow trials are evenly distributed throughout the sequence, even if the two sessions produce identical CoV values. This difference is theoretically meaningful and clinically relevant for the monitoring contexts PRECISE targets.
To address MSSD’s scale dependence, the protocol specifies normalized MSSD computation as:
nMSSD = MSSD / M_RT^2
This normalization scales MSSD by the squared mean RT, producing a dimensionless index that is interpretable across participants and sessions with different mean speed profiles. nMSSD is analogous in its normalization logic to CoV, applied to the successive-difference domain.
nMSSD is computed and logged for all sessions in which 40 or more valid sequential trial pairs are available after data cleaning. The higher minimum for nMSSD than for CoV reflects the fact that successive-difference statistics are inherently noisier per data point than cross-trial statistics, and 40 valid pairs provides a more defensible floor for a reportable estimate.
The lapse-detection threshold for nMSSD is defined at the session level as an nMSSD value more than two standard deviations above the individual’s rolling nMSSD baseline. When this threshold is crossed, the session is flagged as lapse-elevated independently of whether its CoV value crosses the drift threshold. This dual-flag architecture allows the system to distinguish between two qualitatively different patterns of elevated IIV: a session in which the entire distribution has widened, captured by CoV, and a session in which the distribution is otherwise normal but punctuated by clustered attentional failures, captured by nMSSD. Both patterns are meaningful, but they may have different interpretive implications and different appropriate responses.
3.2.3 Primary Metric for Executive Function Tasks: Raw iSD
For executive function tasks including the anti-saccade paradigm, Go/No-Go inhibition tasks, and working memory span, the primary IIV metric is raw iSD rather than CoV. This deviation from the reaction time metric selection requires explicit justification.
Executive task performance reflects a qualitatively different cognitive process than simple or choice reaction time. In the anti-saccade paradigm specifically, trial-to-trial variability in RT reflects not only processing speed fluctuation but also variability in the success and timing of the executive override process, the frontal inhibitory signal that must suppress the reflexive prosaccade and generate a voluntary countermove. Normalizing this variability by mean RT would remove part of the variance of interest, because faster executive override is not simply faster processing in the same sense that faster simple RT is faster sensorimotor conduction. The two processes are neurally and computationally distinct, and the normalization appropriate for one is not necessarily appropriate for the other.
The published normative literature on anti-saccade IIV, including Hutton and Ettinger (2006) and the framework established by Munoz and Everling (2004), uses raw iSD as the variability metric. Maintaining consistency with the reference literature supports the interpretability of PRECISE-ICS estimates relative to established benchmarks and facilitates eventual concurrent validity comparisons.
For working memory span tasks, session-level iSD is computed across the span accuracy scores across trials, where available, or across response times when accuracy is held constant. The interpretation of IIV in span tasks requires additional care because performance can be close to ceiling for some participants on some sessions, compressing the variability estimate in a way that reflects task difficulty rather than cognitive consistency. Sessions in which mean accuracy exceeds 95 percent are flagged as potential ceiling effects, and their iSD estimates are noted as potentially underestimating true variability.
3.3 Data Cleaning Rules
3.3.1 Anticipatory Response Exclusion
Any trial with a response time below 100 milliseconds is classified as an anticipatory response and excluded from all IIV metric computation. This threshold represents the physiological floor for a genuine perceptual-motor reaction. The full sensorimotor pathway from stimulus onset to finger movement requires approximately 70 to 100 milliseconds to complete under optimal conditions: the visual signal must travel from the retina through the lateral geniculate nucleus of the thalamus to primary visual cortex, undergo cortical processing sufficient to identify the stimulus and initiate a response plan, and generate a motor output through the corticospinal tract to the effector. Responses faster than 100 milliseconds cannot represent this full pathway operating genuinely. They reflect button-holding artifacts, touchscreen pressure-registration variability, or input event misfires in the browser’s event queue.
Anticipatory responses are excluded from IIV metric computation but are retained in the session record and their rate is computed as a session-level quality index. A session in which more than 10 percent of administered trials produce anticipatory responses is flagged as a data quality concern. High anticipatory response rates may indicate participant inattention, device input artifacts, or deliberate gaming of the assessment, all of which are interpretively relevant and should be logged for downstream analysis.
3.3.2 Lapse Response Identification and Differential Handling
Lapse responses, defined operationally as reaction times exceeding three standard deviations above the individual’s within-session trimmed mean, require differential handling that depends on which IIV metric is being computed. This differential handling is one of the most methodologically consequential decisions in the protocol and requires careful justification.
For CoV computation, lapse trials are excluded before the iSD and mean are computed. The rationale is that a single extreme lapse can inflate the iSD substantially and would dominate the session-level CoV estimate in a way that misrepresents the individual’s typical variability profile. If a participant produces 29 trials with RT values clustering tightly between 280 and 340 milliseconds and one trial of 1,200 milliseconds due to a momentary distraction, the resulting iSD reflects the 1,200-millisecond outlier far more than it reflects the genuine variability structure of the other 29 trials. CoV computed from the trimmed distribution more accurately characterizes what the monitoring system is designed to measure: the individual’s ongoing moment-to-moment variability in genuine reactive performance.
For nMSSD computation, lapse trials are retained. MSSD’s sensitivity to attentional lapses is precisely the theoretical reason it is included in the protocol as a supplementary metric. Excluding lapse trials before computing MSSD would eliminate the signal the metric is designed to detect. The lapse trial and the adjacent trial pair it is part of are the data. Removing them would produce an nMSSD value that underestimates the attentional disruption of the session.
This means that CoV and nMSSD are computing variability from partially different trial sets in any session that contains lapse trials. This is not a methodological inconsistency. It is a deliberate design choice that reflects the different theoretical questions the two metrics are answering. CoV answers: what is this individual’s typical consistency during genuine reactive performance? nMSSD answers: how frequently and severely did this individual’s attention fail during this session? Both questions are worth answering. They require different data treatments.
The three-SD threshold for lapse classification uses the within-session trimmed mean as its reference rather than the individual’s historical baseline mean, because the protocol cannot assume that baseline data is available at the beginning of the monitoring period, and because within-session mean provides a more appropriate reference for identifying within-session outliers. The trimmed mean used here is the mean of the central 90 percent of within-session trials, excluding the top and bottom 5 percent before computing the SD threshold.

Image Credit: ChatGPT
3.3.3 Timeout Trial Exclusion and Logging
Trials in which no response is recorded within the task’s designated response window, defined as 3,000 milliseconds for reaction time tasks and 5,000 milliseconds for executive function tasks, are excluded from all IIV metric computation and logged as timeout events. Timeout trials differ from lapse trials in that they represent complete attentional disengagement from the task rather than a slow response. Their exclusion from IIV computation is appropriate because including them would conflate missing data with extreme variability, and because a participant who missed a trial due to looking away from the screen has not produced a genuine reaction time at all.
Timeout events are tracked across sessions as a longitudinal quality and engagement indicator. A session in which more than 15 percent of administered trials result in timeouts is flagged as a session-level concern, and the session’s IIV estimates are designated as low-confidence and excluded from baseline updating. Sustained high timeout rates across multiple sessions may indicate engagement problems, environmental disruption, or genuine attentional difficulties, and all three interpretations are worth distinguishing through follow-up assessment.
3.3.4 Minimum Valid Trial Threshold and Session Validity Classification
A session is classified as valid for baseline updating and longitudinal analysis only if it meets the following criteria simultaneously: at least 30 valid trials remain after anticipatory response exclusion and lapse response exclusion for CoV computation; the anticipatory response rate is below 10 percent of administered trials; the timeout rate is below 15 percent of administered trials; and the within-session fatigue index does not exceed 35 percent, as defined and justified in Section 3.2.1.
Sessions that fail one criterion are classified as low-confidence and logged without contributing to baseline updating. Sessions that fail two or more criteria are classified as invalid and excluded from the longitudinal record, though their raw data is retained for data quality audit purposes.
The practical implication of this classification system for task design is direct. If anticipated exclusion rates across all cleaning criteria produce an average loss of 10 to 20 percent of administered trials, then a task must administer at least 36 to 38 trials to reliably produce the 30 valid trials needed for a valid session. This should be treated as a hard design constraint for any task implementation that intends to produce PRECISE-ICS-compliant IIV estimates.
3.4 Baseline Formation
3.4.1 The Minimum Session Requirement: Justification and Limitations
A single session cannot produce a meaningful IIV baseline. Neither can two or three sessions. The session-level CoV estimate for a given individual contains both genuine state variability reflecting the individual’s cognitive status on that day, and measurement noise reflecting the sampling variability inherent in estimating a standard deviation from 30 to 50 trials. For a rolling mean and standard deviation of session-level CoV values to be stable enough to serve as a reference distribution against which future sessions can be evaluated, a sufficient number of sessions must have accumulated that the mean and standard deviation of the session-level CoV estimates have converged to values that are not dominated by sampling variability.
The PRECISE-ICS Standard proposes a minimum of eight sessions for initial baseline formation. This threshold is operationally derived rather than empirically validated, and that distinction is stated explicitly and without apology. The argument for eight sessions has two components.
The first is statistical. For session-level CoV estimates with an assumed true session-to-session coefficient of variation of approximately 15 to 20 percent, which is a reasonable expectation based on the adult IIV literature’s documentation of test-retest reliability in RT variability measures, a sample of eight sessions allows the estimation of a baseline mean with a standard error of approximately 5 to 7 percent of the true mean, and a baseline standard deviation with moderate stability. These are rough estimates that depend heavily on the true within-person variability structure, which has not been characterized in adolescent populations using browser-based assessment. They are offered as order-of-magnitude justification for the threshold, not as precise psychometric guarantees.
The second is ecological. Eight sessions at a cadence of two to three sessions per week corresponds to a baseline formation window of three to four weeks. This is a realistic commitment for adolescent users in school or sports program contexts. A baseline formation requirement of twenty sessions would produce better statistical stability but would lose most participants before baseline was achieved. A requirement of four sessions would be ecologically feasible but statistically premature. Eight represents a considered balance between statistical defensibility and ecological viability, with the explicit acknowledgment that the optimal number is an empirical question that this protocol generates for prospective testing.
For comparative context: ImPACT composite scores have been shown to remain considerably stable across one, two, and three-year test-retest intervals in high school athletes (Schatz et al., 2006), which supports the use of annual baselines as sufficient reference points for concussion management. PRECISE’s proposed requirement of eight valid sessions over three to four weeks is a substantially more granular standard, one designed for a system attempting to track week-to-week and month-to-month cognitive fluctuations rather than post-injury recovery relative to a distant pre-season measurement.
3.4.2 Session Cadence During Baseline Formation and Ongoing Monitoring
During the baseline formation period, the protocol specifies a minimum cadence of three sessions per week for a minimum of three weeks, producing a minimum of nine administered sessions. This exceeds the eight-session statistical minimum by one session deliberately: the extra session provides a buffer for one session-level validity failure, such as a data quality flag or a timeout-heavy session, without requiring the participant to extend the baseline formation window. If all nine sessions are valid, the baseline is computed from all nine. If one session is classified as low-confidence or invalid, the baseline is computed from the remaining eight, meeting the statistical minimum.
Sessions during baseline formation should be completed at varied times of day, not at a fixed daily time, in order to sample the individual’s IIV across their natural circadian performance range. An individual who consistently tests at 8 AM before school will produce a baseline calibrated to their early-morning cognitive state, which may not generalize to their afternoon performance profile. Varying session timing during baseline formation produces a more representative sampling of the individual’s true intra-individual variability range.
During ongoing monitoring after baseline establishment, the protocol proposes a minimum cadence of two sessions per week. This frequency is expected to be sufficient to detect week-to-week state changes associated with events such as concussion, acute illness, major sleep disruption, or examination-period stress, while remaining ecologically feasible as a long-term commitment. Whether two sessions per week is in fact the optimal monitoring frequency for the changes PRECISE is designed to detect is an empirical question addressed in the validation agenda. It also produces sufficient data for the rolling baseline window to remain responsive to genuine longitudinal trends without becoming so data-sparse that the reference distribution degrades.
These cadence specifications are operationally proposed. The optimal monitoring frequency is an empirical function of the time-constant of the cognitive changes PRECISE is designed to detect, which varies across the target monitoring contexts. A concussion produces acute changes measurable within 24 to 48 hours and a recovery trajectory measured in weeks. Chronic fatigue accumulation produces slower drift measured in weeks to months. Academic stress produces changes that may emerge over days and resolve within a week of an examination period ending. These different time-constants suggest that a fixed monitoring cadence will be optimal for some use cases and suboptimal for others, and that adaptive cadence recommendations based on the monitoring context are a logical extension of this protocol that belongs in its next version.
3.4.3 The Rolling Baseline Window: Specification and Rationale
After the initial baseline formation period, the PRECISE-ICS Standard does not use all historical sessions as the baseline reference. Instead, it specifies a rolling baseline window of the most recent twelve sessions, updated after each new valid session.
The rolling window serves two distinct purposes that both matter for the protocol’s scientific defensibility in adolescent populations.
The first purpose is developmental responsiveness. As established in Section 2.3, the expected developmental trajectory of IIV in adolescents is a gradual decline as prefrontal networks mature. If the baseline is computed from all historical sessions including those from six or twelve months ago, the reference distribution will lag the individual’s developmental trajectory. A participant whose baseline CoV was 0.22 at enrollment and has declined to 0.15 after twelve months will trigger drift alerts when their current sessions produce CoV values of 0.17, even though 0.17 is lower than the enrollment baseline. A rolling window that anchors to the most recent twelve sessions will have a reference mean that has tracked downward with that trajectory, reducing this class of false positive.
The second purpose is baseline currency. Sessions from twelve or eighteen months ago should carry less weight in the reference distribution than sessions from last month. A participant who had a high-stress academic semester six months ago and whose IIV was elevated throughout that period should not have their current performance evaluated against a baseline contaminated by that stress-elevated period. The rolling window ensures that the reference distribution represents the individual’s recent stable state rather than an averaged mixture of stable and disrupted periods from their past.
The twelve-session rolling window corresponds to approximately six weeks of monitoring at the recommended two-sessions-per-week cadence. This is a proposed estimate of the timescale over which the reference distribution should be anchored. The relevant validation question is whether twelve sessions provides sufficient stability for the reference distribution while remaining responsive enough to track genuine developmental trends without excessive lag. This is the third primary empirical question the protocol generates, after baseline convergence and rolling window optimization.
Sessions that are classified as low-confidence or invalid according to the criteria in Section 3.3.4 are excluded from the rolling window. This means the rolling window always contains twelve valid sessions, not twelve calendar-time sessions, and that the calendar time span of the rolling window will be longer for participants with higher session invalidity rates.
3.5 Longitudinal Drift Detection
3.5.1 Reference Distribution Computation
Within the rolling baseline window, the protocol computes two quantities that define the individual’s personal IIV reference distribution: the personal CoV baseline mean (CoV_BL) and the personal CoV baseline standard deviation (CoV_BL_SD), computed across all valid sessions in the rolling window.
These two quantities represent the individual’s stable IIV state and the normal session-to-session variability of that state. CoV_BL_SD is critical to the drift detection system because it calibrates the detection threshold to the individual’s own volatility. A participant whose session-level CoV values naturally vary across a wide range from week to week requires a larger absolute change to register as meaningful drift than a participant whose session-level CoV values are tightly clustered. Applying a fixed absolute threshold to both would produce substantially elevated false positive rates for the high-volatility participant and inadequate sensitivity for the low-volatility participant. The individual-referenced SD threshold is proposed to address this by scaling detection sensitivity to the individual’s own natural variability profile, though whether it does so effectively in practice is an empirical question.
3.5.2 Drift Threshold Specification
A new session’s CoV value (CoV_S) is evaluated against the personal reference distribution using the following threshold structure.
Significant positive drift, indicating a session with substantially elevated IIV relative to the individual’s stable baseline, is flagged when:
(CoV_S - CoV_BL) > 2.0 * CoV_BL_SD
Moderate positive drift, warranting elevated monitoring but not high-priority notification, is flagged when:
1.5 * CoV_BL_SD < (CoV_S - CoV_BL) <= 2.0 * CoV_BL_SD
These thresholds are adapted from the Reliable Change Index methodology proposed by Jacobson and Truax (1991), which provides a statistical framework for distinguishing meaningful change from measurement noise in longitudinal assessment contexts. The adaptation to within-person IIV monitoring is a proposal rather than a direct application of the validated RCI procedure. The 2.0 SD threshold corresponds approximately to a change that would occur by chance in fewer than 5 percent of sessions under a normal distribution assumption for the session-level CoV estimates, and the 1.5 SD threshold corresponds to a change occurring by chance in fewer than 7 to 8 percent of sessions. Both percentile interpretations depend on the assumption that session-level CoV values are approximately normally distributed in this population, which has not been tested. If the session-level CoV distribution is substantially skewed or heavy-tailed, the implied false positive rates will diverge from these estimates and distribution-specific thresholds will be needed. Whether the proposed thresholds perform as intended is therefore one of the empirical questions the validation agenda in Section 7 is designed to address.
3.5.3 The Directional Asymmetry of Drift Flagging
The protocol flags only positive drift, meaning increases in IIV relative to baseline, as the primary alert class. Negative drift, meaning decreases in IIV, is tracked and logged but does not generate user-facing notifications. This asymmetry requires explicit justification, because it is a deliberate departure from a symmetric detection scheme.
In adolescent populations, a decrease in IIV below the personal baseline has three plausible interpretations: genuine improvement in neural regulation reflecting effective cognitive training or recovery; normal developmental maturation producing the expected reduction in variability; or a practice effect in which increased familiarity with the task protocol reduces response variability without reflecting genuine changes in underlying cognitive stability. All three of these interpretations are benign or positive, and generating a user-facing alert for a decrease in IIV in this population would create confusion without clinical justification.
This does not mean that large negative drift is ignored. Sessions producing CoV values more than 1.5 standard deviations below the baseline mean are flagged in the system log as potential practice effects or developmental shifts that may warrant baseline window reanchoring. If negative drift persists across the rolling window, it will naturally shift the baseline mean downward, and future alerts will be calibrated to the new lower reference. This is the intended behavior for a developmentally responsive monitoring system, and whether it performs as intended is a question for prospective validation.
3.5.4 The Persistence Rule: Reducing Clinically Unjustified Alarms
A single session crossing the significant drift threshold does not generate a high-priority user notification. The PRECISE-ICS Standard specifies a persistence rule requiring that significant positive drift be sustained across a minimum of three consecutive valid sessions before a high-priority notification is issued.
The rationale is that a single elevated CoV session has many plausible benign explanations: acute sleep deprivation the preceding night, environmental distraction during the session, a stressful day at school, or simple sampling variability producing an above-baseline estimate. A single elevated session is, in other words, a relatively weak signal that is consistent with many causes, most of which are transient and require no intervention. Three consecutive sessions crossing the significant drift threshold is substantially more informative: it represents a sustained pattern that is unlikely to reflect transient causes and is much more consistent with a genuine state change that warrants attention.
The persistence rule represents a deliberate tradeoff between sensitivity and false positive control. It sacrifices the ability to flag acute single-session events, including acute concussive events in which IIV may spike dramatically in a single session, in favor of dramatically reduced false positive rates for sustained monitoring. For PRECISE’s current deployment context, which targets general cognitive health monitoring in adolescent populations rather than acute injury detection specifically, this paper argues the tradeoff is appropriate, though prospective validation should test whether the three-session threshold optimally balances sensitivity and specificity in this population. A future version of the protocol designed specifically for post-concussive monitoring would need to revisit this rule, potentially adding a single-session high-threshold alert alongside the three-session persistence rule to capture acute events without inflating the chronic false positive rate.

Image Credit: ChatGPT
4. Adolescent-Specific Protocol Adaptations
4.1 Age-Stratified Baseline Window Recommendation
The expected rate of developmental IIV change is higher in early adolescence than in late adolescence, reflecting the faster rate of prefrontal maturation and synaptic pruning in the earlier years of the second decade. A rolling baseline window calibrated for a 17-year-old whose developmental IIV trajectory is relatively stable may be too long for a 13-year-old whose IIV is expected to be declining faster as a result of more active ongoing cortical reorganization.
The protocol proposes, as a provisional recommendation pending empirical validation, an age-stratified window specification: a rolling window of eight sessions for users aged 12 to 14, and the standard twelve-session window for users aged 15 to 19. The shorter window for younger adolescents ensures that the reference distribution tracks more closely with the faster developmental trajectory, reducing the risk of the baseline lagging the developmental change and generating false positive drift alerts.
This recommendation is based on the developmental neuroscience literature’s characterization of the relative rates of prefrontal maturation across adolescent age bands. It is not derived from empirical calibration of rolling window lengths in this population, because such data does not exist. Prospective validation should explicitly test whether different window lengths produce different false positive rates across age bands, and should use this empirical data to refine the age-stratified recommendation.
4.2 Interpreting IIV Trajectories Against Developmental Background
The most important practical challenge in adolescent IIV monitoring is distinguishing IIV changes that reflect genuine cognitive state changes from IIV changes that reflect normal developmental maturation. The rolling baseline window addresses this problem structurally, by ensuring that the reference distribution tracks with the developmental trajectory rather than anchoring to a distant historical baseline. But the challenge does not disappear entirely, because not all participants will show the expected developmental IIV reduction, and some will show irregular trajectories that are difficult to interpret without additional context.
The protocol therefore specifies an annual developmental calibration procedure for users monitored continuously across twelve or more months. At each twelve-month mark, the system computes the mean session-level CoV for the most recent three-month block and compares it to the mean for the three-month block twelve months earlier. A decrease of more than 15 percent of the baseline mean CoV value is classified as a developmental reduction and logged as a developmental milestone rather than a monitoring event. This prevents the system from generating drift alerts in response to the natural long-term developmental IIV reduction that is expected in this age group.
A failure to show any IIV reduction across twelve months of continuous monitoring, in a participant who showed no significant drift events during that period, is itself an interpretively interesting finding. It may reflect a participant whose development is proceeding on a slower timeline within the normal range, or it may reflect a genuinely flat developmental trajectory that warrants further investigation. The protocol does not prescribe an action for this case, because the appropriate response depends on context that the system cannot assess autonomously. It logs the finding and surfaces it for review.
4.3 Comorbidity and Context Logging
Adolescent IIV is sensitive to a wide range of factors that are part of normal adolescent life: sleep variability, academic stress, physical illness, menstrual cycle phase in female participants, and emotional stress. These factors are not pathological, but they are relevant to IIV interpretation. A monitoring system that detects an IIV elevation without any contextual information about whether it occurred during examination week, following a sleep-disrupted night, or during a week in which the participant reported high stress, cannot distinguish state-appropriate variability from monitoring-relevant change.
The PRECISE-ICS Standard recommends the implementation of a brief session-level context log in which participants record, at the time of each session, their estimated hours of sleep the preceding night, their self-rated stress level on a three-point scale, and whether they have experienced any recent illness, physical injury, or significant life event. This context log does not constitute a validated clinical assessment instrument, and its data should not be used for clinical interpretation. Its purpose is to enable post-hoc stratified analysis of IIV data: the monitoring system can identify sessions that occurred in high-stress or sleep-deprived contexts and can present IIV trends with and without these sessions included, providing the user and any reviewing clinician with a more nuanced picture of the longitudinal pattern.
5. Statistical and Psychometric Considerations
5.1 The Reliability of Browser-Based CoV Estimates
The reliability of a session-level CoV estimate, meaning the degree to which it would produce a similar value if the same participant were tested again under similar conditions, is a function of both the number of trials contributing to the estimate and the true within-person session-to-session variability. In the adult RT variability literature, test-retest reliability for CoV-like measures over short intervals has been documented at acceptable levels for research purposes in laboratory settings, typically in the range of 0.65 to 0.80 depending on the number of trials, the task type, and the retest interval.
Comparable reliability data for browser-based CoV estimates in adolescent populations do not exist in the published literature. This is an honest gap that the prospective validation research agenda must address. Until reliability coefficients are available for browser-based adolescent CoV estimates specifically, the protocol’s 30-trial minimum and 8-session baseline formation requirement should be understood as reasonable operational constraints for a monitoring system operating without this data, rather than as thresholds derived from a known reliability function.
5.2 The Signal-to-Noise Structure of the Longitudinal Record
The PRECISE-ICS longitudinal record accumulates three sources of variation in session-level CoV estimates over time. The first is genuine between-session state variability, meaning real fluctuations in the participant’s cognitive consistency that reflect factors such as sleep, stress, and health status. This is the signal the monitoring system is designed to detect.
The second is measurement noise, meaning the sampling variability inherent in estimating a standard deviation from 30 to 50 trials. This source inflates the apparent between-session variability of CoV estimates relative to the true between-session variability, making drift detection more difficult.
The third is browser-based timing artifact, meaning the variable latency introduced by the browser environment that inflates iSD without reflecting genuine cognitive variability. The cleaning procedures in Section 3.3 and the timing architecture in Section 6 are designed to minimize this source.
The detection thresholds specified in Section 3.5.2 are calibrated to the total observed between-session variability of session-level CoV estimates, which includes all three sources. This means the detection system is calibrated to the noise floor as observed rather than to the true signal magnitude. The consequence is that the 2.0 SD threshold implies different absolute CoV changes for different participants, depending on how much measurement noise is present in their individual record. A participant with very tight session-to-session CoV estimates will trigger the drift alert at a smaller absolute change than a participant with high natural session-to-session variability, which is the intended behavior for a system designed to detect changes that are unusual for that individual rather than unusual in absolute terms.
5.3 Multiple Comparisons and the Cost of Continuous Monitoring
A monitoring system that evaluates each new session against a drift threshold is implicitly conducting a statistical test each time a session is completed. Over a monitoring period of fifty sessions, the cumulative probability of at least one false positive alert, even at a per-session false positive rate of 5 percent, is approximately 92 percent under independence assumptions. The persistence rule reduces this substantially, because it requires three consecutive false positives to trigger a notification, but it does not eliminate the multiple comparisons problem entirely.
The protocol acknowledges this issue directly and makes no claim that the drift detection system is operating at a 5 percent family-wise error rate across the full monitoring period. The 2.0 SD threshold is a per-session threshold, and its cumulative implications depend on the dependence structure of the session-level CoV estimates across time, which has not been characterized in this population.
The practical mitigation for this issue is the contextual interpretation guidance built into the notification system: all drift alerts generated by the system are accompanied by explicit statements that a single alert, or even a three-session alert, does not constitute a clinical conclusion, and that the appropriate response is continued monitoring and contextual interpretation rather than immediate action. Reducing the per-session threshold to produce a lower cumulative false positive rate would, this paper proposes, sacrifice sensitivity to genuine events in a way that reduces the monitoring system’s clinical utility, though the empirical tradeoff at different threshold levels is a question the validation agenda should address.
6. Browser-Based Timing: A Full Treatment of Limitations and Mitigations
6.1 The Sources of Timing Noise in Browser Environments
Laboratory-grade cognitive assessment achieves its timing precision through a set of hardware and software controls that browser-based systems cannot replicate. Dedicated testing hardware uses interrupt-driven response capture, where the participant’s button press triggers a hardware interrupt that timestamps the response at the level of the operating system kernel, bypassing the browser’s event queue entirely. Display hardware is synchronized to known refresh rates with explicit frame-timing guarantees. The testing environment is controlled for every potential source of variability.
Browser-based assessment operates under fundamentally different constraints. The JavaScript runtime that executes the task logic shares a CPU event loop with browser rendering processes, plugin processes, and potentially other applications running on the same device. Timer APIs operate at the level of the JavaScript runtime rather than the operating system kernel, introducing a layer of indirection between the cognitive event of interest, which is the participant’s response, and the timestamp that records it. Display rendering is asynchronous: the JavaScript instruction that changes the stimulus display is executed immediately, but the visual update queues for the next rendering frame, which on a 60 Hz display occurs approximately every 16 milliseconds. Privacy-motivated timer coarsening, introduced in browser implementations in response to the Spectre and Meltdown speculative execution vulnerability class, further reduces the effective resolution of some timer APIs below their theoretical specification.
These factors produce two distinct categories of timing error. Constant offset error is the systematic difference between the time a stimulus is logically specified and the time it is physically rendered on screen. This error is introduced by the rendering queue latency and produces a systematic inflation of all RT estimates without increasing their variability. It is a threat to absolute performance benchmarking but not to IIV measurement, because a constant offset shifts the mean without widening the distribution. Variable latency error is the session-to-session and trial-to-trial variation in rendering and input event timing caused by CPU load variability, garbage collection pauses, and device-specific hardware characteristics. This category of error directly inflates iSD and therefore CoV without reflecting any genuine change in the participant’s cognitive consistency, and it represents the more serious threat to IIV measurement validity.
6.2 Mitigations Implemented in PRECISE
The PRECISE system addresses the timing precision problem through an architecture that should be understood as the minimum standard for any implementation claiming to produce PRECISE-ICS-compliant IIV estimates.
The high-resolution timer API performance.now() is used for all reaction time measurements in place of Date.now(). The performance.now() API provides sub-millisecond theoretical resolution and is not subject to the same privacy-motivated coarsening as wall-clock APIs in most current browser implementations, though this distinction may change in future browser versions as the security landscape evolves.
Stimulus onset timestamps are recorded inside requestAnimationFrame callbacks rather than at the point of the JavaScript instruction that initiates the stimulus display change. The requestAnimationFrame callback fires at the moment the browser is preparing to paint the next frame, synchronizing the timestamp to the actual display update rather than to the execution queue. This eliminates the systematic rendering-queue offset that would otherwise be present between the logical and physical stimulus onset.
Response capture uses pointerdown events rather than click events. The click event fires after both the initial contact and the release of the input, which introduces 30 to 80 milliseconds of additional latency depending on how long the participant holds the button. pointerdown fires at initial contact, which is the correct cognitive event: the moment the neural signal has traveled from the decision stage to the effector. Using click events for RT measurement in a browser-based cognitive paradigm introduces a systematic latency artifact whose magnitude varies with participant button-pressing behavior, which would inflate both mean RT and session-level iSD in ways that are impossible to correct post-hoc.
Inter-trial intervals are jitter-randomized across a range of 1,000 to 3,000 milliseconds using a pseudorandom sequence. Fixed inter-trial intervals introduce temporal predictability that allows participants to anticipate stimulus onset and begin their motor preparation early, producing artificially fast and artificially consistent responses that underestimate genuine IIV. Jitter-randomized intervals eliminate this anticipatory preparation and ensure that each trial’s response reflects genuine reactive processing.

Image Credit: ChatGPT
6.3 The Residual Noise Floor and Its Implications for IIV Measurement
After applying the mitigations described in Section 6.2, the residual timing noise in the PRECISE implementation is estimated at 5 to 15 milliseconds under optimal conditions, based on published assessments of similar browser-based RT implementations using equivalent API choices. This estimate carries substantial uncertainty because device-level variability, particularly across the heterogeneous range of hardware on which browser-based assessment runs, can produce timing noise substantially above this range on low-powered or thermally throttled devices.
The critical question for IIV validity is whether this residual noise floor is small relative to the IIV thresholds used in drift detection. This paper argues that it is, for the following reason, though this argument requires prospective empirical confirmation. The drift detection threshold of 2.0 standard deviations above the individual’s baseline CoV_BL_SD represents a change in session-level CoV that is typically an absolute change of 0.04 to 0.08 CoV units for participants whose natural session-to-session variability is in the range documented in the adult RT literature. For a participant with a mean RT of 300 milliseconds and a session-to-session CoV standard deviation of 0.03, a 2.0 SD alert corresponds to a CoV change of 0.06, which corresponds to an iSD change of approximately 18 milliseconds on top of a stable mean. A timing noise floor of 5 to 15 milliseconds is not trivially small relative to this threshold, but it is proposed to be substantially below it when the natural averaging across 30 or more trials is considered. The session-level CoV estimate is an average-based statistic: constant timing noise adds a fixed amount to every trial’s RT and therefore adds very little to the iSD; and random variable timing noise of 5 to 15 milliseconds, when averaged across 30 trials, contributes a standard error of approximately 3 milliseconds to the session-level iSD estimate. If this reasoning holds empirically, the contribution is small enough relative to the drift detection threshold that the protocol remains defensible as a longitudinal monitoring instrument. Validation Question 1 in Section 7 is designed in part to test this assumption directly.
6.4 Device Calibration as a Future Mitigation
The protocol specifies pre-session device calibration as a recommended future extension that is not currently implemented in the PRECISE system but should be a development priority before any institutional deployment of the IIV monitoring protocol. Device calibration involves presenting stimuli at known intervals at the start of each session, measuring the device’s rendering latency through the difference between expected and observed timing, and subtracting the estimated device-specific offset from subsequent trial timestamps.
This approach is already implemented in the PRECISE speech assessment module for Web Speech API transcription delay correction, demonstrating that the architectural approach is feasible within the browser-based framework. Extending it to the visual task timing system would reduce the device-level component of variable latency and improve the session-to-session stability of CoV estimates on heterogeneous hardware. The expected improvement in CoV estimate stability is difficult to quantify without empirical data, but reducing the residual noise floor from the current 5 to 15 millisecond range to a post-calibration range of 2 to 8 milliseconds would meaningfully improve the signal-to-noise ratio of the longitudinal monitoring record.
7. Proposed Validation Research Agenda
7.1 The Minimum Validation Burden
A methodology proposal without a validation agenda is a methodology proposal without a path to scientific credibility. The following research questions constitute the minimum empirical program required to establish the PRECISE-ICS Standard’s defensibility as a monitoring instrument. They are ordered by logical priority: later questions cannot be meaningfully answered without first answering earlier ones.
7.2 Primary Validation Questions
Question 1: Session-level CoV reliability in adolescent populations using browser-based assessment. Before the protocol’s minimum trial and session requirements can be empirically justified, the test-retest reliability of session-level CoV estimates in this specific population and deployment context must be characterized. This requires a sample of at least 40 adolescent participants completing two sessions within a 72-hour window under comparable conditions, computing session-level CoV from both sessions, and computing the intraclass correlation coefficient between the two estimates. If ICC is below 0.65, the 30-trial minimum is insufficient and must be increased. If ICC is above 0.80, the minimum may be reduced without sacrificing reliability. This study should be designed and executed before any institutional adoption of the protocol.
Question 2: Baseline convergence as a function of session count. At what session count do within-person CoV estimates converge to within an acceptable tolerance of their longer-term mean in adolescent populations? This requires a sample of at least 30 adolescent participants completing a minimum of twenty sessions each, with baseline stability analyzed as the session count required for the rolling mean CoV to stabilize within 10 percent of the twenty-session mean. If the distribution of convergence session counts indicates that most participants stabilize by session eight, the protocol’s eight-session minimum is validated. If most participants require twelve or more sessions, the minimum must be revised.
Question 3: Rolling window optimization. What window length minimizes false positive drift detections due to developmental IIV reduction while retaining sensitivity to genuine state-based IIV increases? This requires either a sample with known experimental state manipulations, such as controlled acute sleep restriction or examination-period stress, or a natural history sample including documented concussion events, allowing the comparison of detection rates and false positive rates across window lengths of six, eight, twelve, and sixteen sessions.
Question 4: Concurrent validity with gold-standard instruments. Do PRECISE-ICS session-level CoV and nMSSD values correlate significantly with simultaneously administered gold-standard assessments? This requires a concurrent validity study in which participants complete both PRECISE assessments and at least one validated instrument, such as Cogstate’s Detection or Identification tasks or a neuropsychologist-administered simple reaction time battery, within the same testing window. The expected correlation should be moderate rather than high, because the instruments measure overlapping but not identical constructs under different testing conditions. A correlation below 0.40 would suggest the measures are not converging on the same underlying construct and would warrant reassessment of the task design.
Question 5: Sensitivity to ecologically relevant state changes. Does the protocol detect IIV changes associated with (a) documented concussion events in student athletes, identified through physician or athletic trainer confirmation; (b) acute sleep restriction of four or fewer hours the preceding night, confirmed through actigraphy or self-report; and © elevated psychological stress during examination periods, confirmed through validated self-report instruments such as the Perceived Stress Scale? Each of these requires a separate study design with appropriate control conditions, and together they constitute the ecological validity core of the validation agenda.
8. Ethics and Data Governance
8.1 Informed Consent, Assent, and Guardian Authorization
Cognitive monitoring in adolescent populations requires age-appropriate assent from participants and informed consent from guardians where required by local regulation. The assent and consent procedures must explicitly communicate the non-diagnostic nature of PRECISE-ICS outputs, the specific data collected and how it is processed, the absence of data transmission in the current implementation, and the right to withdraw without consequence. Institutional research ethics board approval should be obtained before any implementation of this protocol in a research context. Applied implementations in school or sports program settings should follow the institutional data governance requirements of the deploying organization.
8.2 Data Minimization and Local Processing
IIV computation requires only timestamped response data: reaction time in milliseconds, response correctness indicator, trial number, and session timestamp. It does not require storage of audio recordings, video data, or any biometric signal beyond the response timing and accuracy measures generated by the cognitive tasks. The PRECISE system’s current architecture computes all metrics locally within the participant’s browser and exports only structured numeric CSVs, which is consistent with data minimization principles appropriate for a research prototype operating without formal institutional oversight.
Future extensions incorporating physiological signals such as heart rate variability should apply the same minimization principle: only derived metrics, not raw waveforms, should be retained unless a specific research protocol with appropriate ethics oversight justifies raw signal storage.
8.3 Interpretation Responsibility and Output Framing
The PRECISE-ICS Standard does not produce diagnostic outputs. A drift alert generated by the system is not a diagnosis of concussion, attention disorder, cognitive decline, or any other condition. It is a statistical signal indicating that the individual’s recent IIV profile departs from their personal reference distribution in a way that may warrant continued monitoring or contextual interpretation. The appropriate response to a drift alert is not clinical action. It is continued monitoring, contextual review, and, if the pattern persists and the context warrants it, consultation with a qualified professional who can interpret the data appropriately.
Every output generated by any system implementing the PRECISE-ICS Standard must communicate this framing clearly, consistently, and in language appropriate for the age and context of the user. Outputs that are framed as diagnostic conclusions, risk scores, or clinical recommendations violate the intent of this protocol and introduce potential harms that the protocol’s scientific design does not justify.
9. Conclusion
9.1 What This Paper Has Done
This paper has proposed a formal, explicit, and testable methodology for computing, interpreting, and longitudinally monitoring intra-individual variability in adolescent cognitive performance using browser-based ecological assessment. Every methodological choice in the protocol has been justified by reference to the comparative IIV literature, the developmental neuroscience of adolescent brain maturation, the psychometric requirements for longitudinal stability, and the specific engineering constraints of browser-based deployment. Every proposed threshold that is operationally derived rather than empirically validated has been identified as such.
The choices made in this protocol are not the only defensible choices. A different research team might reasonably prefer a shorter rolling baseline window, a different drift threshold, or a different minimum trial count. What this protocol provides is not the definitive answer to those questions. It provides an explicit, arguable starting point from which those questions can be empirically tested and refined.
9.2 What This Paper Has Not Done
This paper has not produced a validated clinical protocol. It has not demonstrated that the proposed thresholds correctly classify participants in empirical data. It has not shown that the rolling baseline window optimally tracks developmental IIV trajectories. It has not established the test-retest reliability of browser-based CoV estimates in adolescent populations. All of these things remain to be done, and the validation research agenda in Section 7 is the roadmap for doing them.
This is not a weakness to be concealed. It is the honest state of the field, which currently has no standardized protocol for this monitoring context at all. A rigorous proposal with explicit limitations and a clear validation agenda is more valuable than a vague claim that is difficult to test or refute. The scientific value of this paper is precisely in making the methodology testable: if the eight-session minimum is wrong, a prospective study will show that. If the rolling window length is wrong, empirical data will demonstrate the correct value. If the drift thresholds produce elevated false positive rates, calibration will correct them. None of that can happen until the methodology is stated explicitly enough to be tested.
9.3 The Broader Argument
The most important claim this paper makes is not technical. It is conceptual.
The cognitive monitoring field has, for decades, organized its assessment philosophy around mean performance as the primary unit of analysis, population norms as the reference standard, and episodic assessment as the measurement model. Those choices produce assessments that are useful for their designed purposes: diagnosing conditions in clinical populations, tracking treatment response in adult patient groups, and providing pre-season concussion baselines for return-to-play decisions. What they do not produce is a monitoring system that can answer the question: is this particular brain behaving differently than it was behaving last week?
That question requires a different philosophy. It requires longitudinal measurement at sufficient frequency to detect week-to-week changes. It requires within-person baselines rather than population norms, because population norms cannot distinguish a person’s stable state from a deviation from it. It requires variability as a primary signal rather than a nuisance to be averaged away, because variability is the feature of the data most sensitive to the neural disruptions the monitoring is designed to detect. And it requires an adolescent-specific framework, because the adolescent brain is not a small adult brain. Its IIV profile is qualitatively different, developmentally dynamic, and interpretable only against the background of what a maturing brain is expected to do.
The PRECISE-ICS Standard is an attempt to provide that framework. It will be wrong in some of its specifics. The validation research agenda will find which specifics and correct them. What this paper is confident about is the underlying argument: that IIV monitoring in adolescent populations is a scientifically grounded, methodologically tractable, and clinically meaningful endeavor that the field currently lacks the protocol infrastructure to pursue systematically. This paper proposes that infrastructure, in enough detail to be tested, and in full acknowledgment of how much remains to be learned.
About Me and About PRECISE
I am Aliza Samir Khoja, a fourteen-year-old innovator at The Knowledge Society (TKS). My focus area is neurotechnology and concussions, with a specific interest in proactive cognitive health monitoring and the translation of laboratory neuroscience into accessible tools.

PRECISE (Predictive Real-time Evaluation of Cognitive Instability using Speech and Executive-functioning) is a staged neurotechnology platform I am building to shift brain health from reactive to dynamic. It is designed for youth, athletes, students, and anyone who wants to understand how their brain is doing across time, without requiring clinical infrastructure.
The roadmap: PRECISE Mini (live) → PRECISE Pro (live) → PRECISE- ICS (article released) → PRECISE Max (in design)
This paper was prepared as part of the PRECISE neurotechnology research program under The Knowledge Society’s Focus framework. All proposed thresholds are operational estimates pending prospective validation. This document does not constitute a medical protocol or clinical guideline. Feedback, critique, and collaboration inquiries are welcomed. If you are researching, building, or working at the intersection of cognitive health, neurotech, and youth performance, I want to hear from you. Find me on LinkedIn or Substack.
References
Beattie, G., & Bradbury, C. (1979). The pause and speech. Linguistics, 17(5–6), 347–360.
Crawford, T. J., Higham, S., Renvoize, T., Patel, J., Dale, M., Suriya, A., & Tetley, S. (2011). Antisaccade performance in normal and clinical populations: Normative data and diagnostic implications. Neuropsychology, 25(1), 1–10. https://doi.org/10.1037/a0020935
Duchek, J. M., Balota, D. A., Tse, C. S., Holtzman, D. M., Fagan, A. M., & Goate, A. M. (2009). The utility of intraindividual variability in selective attention tasks as an early marker for Alzheimer’s disease. Neuropsychology, 23(6), 746–758. https://doi.org/10.1037/a0016583
Hultsch, D. F., Strauss, E., Hunter, M. A., & MacDonald, S. W. S. (2002). Intraindividual variability, cognition, and aging: A longitudinal study. Neuropsychology, 16(2), 199–207. https://doi.org/10.1037/0894-4105.16.2.199
Hutton, S. B., & Ettinger, U. (2006). The antisaccade task as a research tool in psychopathology: A critical review. Psychophysiology, 43(3), 302–313. https://doi.org/10.1111/j.1469-8986.2006.00403.x
Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19. https://doi.org/10.1037/0022-006X.59.1.12
Luna, B., Garver, K. E., Urban, T. A., Lazar, N. A., & Sweeney, J. A. (2004). Maturation of cognitive processes from late childhood to adulthood. Child Development, 75(5), 1357–1372. https://doi.org/10.1111/j.1467-8624.2004.00745.x
MacDonald, S. W. S., Hultsch, D. F., & Dixon, R. A. (2003). Performance variability is related to change in cognition: Evidence from the Victoria Longitudinal Study. Psychology and Aging, 18(3), 510–523. https://doi.org/10.1037/0882-7974.18.3.510
Mullen, J., Hurd, M., Hall, B., & Bhattacharya, M. (2014). Effect of sport-related concussion on clinically measured simple reaction time. British Journal of Sports Medicine, 48(2), 112–118. https://doi.org/10.1136/bjsports-2012-091579
Muñoz, D. P., & Everling, S. (2004). Look away: The anti-saccade task and the voluntary control of eye movement. Nature Reviews Neuroscience, 5(3), 218–228. https://doi.org/10.1038/nrn1345
Schatz, P., Pardini, J. E., Lovell, M. R., Collins, M. W., & Podell, K. (2006). Sensitivity and specificity of the ImPACT test battery for concussion in athletes. Archives of Clinical Neuropsychology, 21(1), 91–99. https://doi.org/10.1016/j.acn.2005.08.001
Tamnes, C. K., Østby, Y., Fjell, A. M., Westlye, L. T., Due-Tønnessen, P., & Walhovd, K. B. (2010). Brain maturation in adolescence and young adulthood: Regional age-related changes in cortical thickness and white matter volume and microstructure. Cerebral Cortex, 20(3), 534–548. https://doi.org/10.1093/cercor/bhp118
Tamnes, C. K., Herting, M. M., Goddings, A. L., Meuwese, R., Blakemore, S. J., Dahl, R. E., & Mills, K. L. (2017). Development of the cerebral cortex across adolescence: A multisample study of inter-related longitudinal changes in cortical volume, surface area, and thickness. Journal of Neuroscience, 37(12), 3402–3412. https://doi.org/10.1523/JNEUROSCI.3302-16.2017
메타데이터
- post_id
- dbbf5ec95899
- slug
- proposed-protocol-for-iiv-computation-in-adolescent-cognitive-monitoring-the-precise-ics-standard-dbbf5ec95899
- url
- https://medium.com/@khoja.aliza/proposed-protocol-for-iiv-computation-in-adolescent-cognitive-monitoring-the-precise-ics-standard-dbbf5ec95899
- canonical_url
- https://medium.com/@khoja.aliza/proposed-protocol-for-iiv-computation-in-adolescent-cognitive-monitoring-the-precise-ics-standard-dbbf5ec95899
- author_url
- https://medium.com/@khoja.aliza
- status
- ok
- fetched_at
- 2026-07-10 08:43:10