Access to and use of mobile phones are increasingly necessary for participation in modern society and for accessing services in health, education, and finance. Despite rising phone access, a persistent digital gender gap exists in India: nationally reported ownership and internet access are higher for men than for women, and this gap is wider in South Asia and in rural areas. Accurate measurement of phone ownership, access, and digital practices is essential to monitor trends and design interventions, but existing survey items vary in conceptualization and often lack testing for local linguistic and cultural contexts.
To develop a robust set of survey questions on digital access and use for a population-level survey in Uttar Pradesh, the research team conducted cognitive interviews to test and improve draft questions. The study reports findings from rural districts of Uttar Pradesh, a primarily Hindi-speaking state, and aims to produce questions that are comprehensible and contextually appropriate for local respondents. Additional cognitive interviewing was planned in Kenya (Kiswahili) and Nigeria (Hausa) to broaden the question library.
The team employed cognitive interviewing: administering draft quantitative survey questions verbatim, eliciting responses using the available response options, then following with qualitative probes to explore respondents’ interpretation and reasoning. Typical testing of each question took one to five minutes: under one minute for the quantitative item and several minutes for probing and discussion. Daily team debriefs after each field day were central to identifying comprehension problems and informing revisions. Questions were revised iteratively and retested across rounds until comprehension improved.
The study identified 118 priority questions by mapping items from major global surveys and adding team-developed questions to fill measurement gaps. Domains included phone ownership, access and use; gendered differences in phone and internet use; norms and attitudes; and concerns or harms related to phone and internet use. A week-long workshop produced Hindi translations through collaborative translation, back-translation, and group discussion to preserve intended constructs before field testing. The 118 questions were split into three interview guides (~40 questions each) to keep interviews under 90 minutes.
Researchers purposively recruited adults from rural Badaun and Jaunpur districts who had fewer years of formal education and who had used a mobile phone at least once in the prior two weeks. The selection aimed to surface comprehension problems that might be masked in more educated or urban samples. As testing progressed, the team increased focus on smartphone users because several items addressed internet use. Guided by thematic saturation, the study completed three rounds of testing with a total of 101 respondents between 01/04/2023 and 27/04/2023.
Seven experienced qualitative researchers (five female, two male), all with master’s-level social science education and regional experience, conducted interviews in pairs (one lead interviewer and one note-taker). Interviews averaged 52 minutes. A logistics manager worked with village leaders and community health workers to recruit participants. All respondents provided oral informed consent for participation and audio recording except one who declined recording.
The analysis began during fieldwork with extensive debrief sessions every eight to 12 interviews. Revisions after debriefs included refining translations, adjusting question order, changing response options, creating alternative wordings, or eliminating questions that did not resonate. Seventy of the most substantive interviews were audio-recorded and translated into English (with key Hindi terms retained) and coded in qualitative analysis software (Dedoose). The remaining 31 interviews were analyzed using detailed debrief notes. Three rounds of testing produced revised versions (version 2 and version 3); a final version 4 was created but not tested further.
Analysis identified seven recurring categories of cognitive and contextual problems that affected respondents’ comprehension and the validity of responses:
Inappropriate terminology: some words or phrases did not match local usage.
Overly complex wording: long or technical phrasing impeded understanding.
Low resonance of digital concepts: terms like personal data, online tracking, privacy policies, and hacking were poorly understood or lacked relevance.
Confusion around permission and supervision: phone sharing and the overlap between permission/supervision and help/support led to ambiguous answers.
Problematic question structures and response options: Likert scales and certain response formats were confusing.
Self-practice bias: respondents’ reported behaviors sometimes reflected aspirational or observed practices rather than actual personal practice.
Unclear time frames and recall expectations: respondents were uncertain about the period to which questions referred.
Many draft questions were found to be incomprehensible to respondents despite careful translation and required revision.
Daily debriefs and iterative testing led to multiple refinements in wording, translation choices, question order, and response options. The testing process resulted in revised question versions that addressed identified mismatches. The tested and final question sets in both Hindi and English are reported in the study’s supplementary materials (S1 Appendix). The authors emphasize that nearly all questions required some revision after cognitive testing.
The study received ethical approval from the Sigma Institutional Review Board in India (IRB Number: 10123/IRB/22–23) and the Johns Hopkins Bloomberg School of Public Health IRB (23938). All participants provided informed oral consent; one participant refused audio recording. Qualitative transcripts contain identifiable household and relationship details; data access is controlled and requests can be made to the Johns Hopkins IRB office as described in the article.
The findings demonstrate the critical role of cognitive interviewing and rigorous translation for improving the measurement of digital access in linguistically and culturally diverse settings. Standardized or globally borrowed survey items may fail to capture locally meaningful constructs, especially for evolving digital concepts. The study’s iterative approach improved question comprehension for a rural, less-educated population and informs plans for broader testing in other languages and regions to build a validated library of survey items for LMIC contexts.