Common mental health conditions, such as depression, anxiety, or other stress-related disorders, constitute approximately 40% of work-related illness cases, and these conditions are mostly manageable and avoidable with appropriate treatment and prevention [,]. On a global scale, around 12 billion workdays are forfeited annually due to the impact of depression and anxiety, resulting in an economic cost of approximately US $1 trillion per year attributed to reduced productivity []. In Europe, productivity loss due to mental health conditions has an estimated annual cost of €400 billion (≈US $460.4 billion) [,]. This productivity loss is a consequence of working with limitations due to an illness (ie, presenteeism) and absence from work due to the illness (ie, absenteeism) [].
Occupational mental health interventions have been increasingly delivered through digital platforms, described as being more flexible, anonymous, and convenient []. This has the potential to increase the accessibility of evidence-based interventions and to reduce costs compared to face-to-face interventions []. There is evidence for moderate effects of occupational e-mental health (OeMH) interventions to reduce stress, insomnia, and burnout, and small treatment effects on depression, anxiety, well-being, and mindfulness [-]. Results regarding work-related outcomes, such as presenteeism and absenteeism, are scarce and inconclusive [,]. In general, there is a high heterogeneity between studies on OeMH interventions in terms of the type of intervention delivered, outcomes, and the quality of their methodology [,]. Furthermore, these mixed results underscore the complexity of assessing the outcomes of digital interventions and indicate that organizational culture, intervention design, and implementation serve as critical mediating factors.
In fact, the implementation of e-mental health interventions in the workplace has some important barriers. Challenges include adapting to new and changing regulatory frameworks, ensuring compatibility with existing systems, addressing employee privacy and security concerns, the rapid change of the digital ecosystem, and managing costs []. Additionally, maintaining treatment adherence and preventing attrition are also recognized as challenges. e-Mental health interventions have usually been associated with high attrition rates []. A recent scoping review indicates that potential successful implementation strategies for OeMH interventions include providing incentives to participants, sending reminders, providing support to users during the intervention, conducting educational meetings, and the involvement of senior management []. The scoping review also concludes that there is an urgent need to enhance the reporting of implementation strategies and to integrate common implementation frameworks with more technology-focused ones to comprehensively address the complexities of e-health implementations [].
Several factors can potentially affect the effectiveness of OeMH interventions, such as the mode of delivery, the content, the approach (eg, tailored and universal), duration of intervention, the use of interventions in combination, and the characteristics of the target population []. Digital interventions that are tailored to users’ needs tend to foster higher engagement []. On the other hand, stigma against individuals with mental illness can deter participation in such interventions []. Finally, the cultural adaptation of workplace interventions is a critical factor for successful implementation, with more favorable outcomes for linguistically and culturally adapted interventions [] and those that consider organizational cultural elements []. Culturally adapted e-mental health interventions seem to have a moderately higher effect in reducing depressive and anxiety symptoms when compared to interventions that are not culturally adapted [].
While occupational mental health interventions delivered via digital platforms have emerged as promising tools for addressing workplace mental health issues, these interventions often lack contextual adaptation to workplace-specific needs and exhibit high variability in outcomes due to differences in design, content, and target populations. A systematic review by Moe-Byrne et al [] highlighted this variability in the effectiveness of digital workplace mental health interventions and emphasized the need for context-sensitive, tailored approaches. Additionally, although digital mental health interventions are growing, many existing programs do not sufficiently consider organizational cultural factors or provide tailored support for small to medium enterprises (SMEs) [].
The EMPOWER [The European Platform to Promote Wellbeing and Health in the Workplace] platform is a multimodal eHealth intervention delivered through a website and an application that was designed to address mental health issues in the workplace and enhance employees’ well-being []. It follows a 3-tiered intervention structure: universal (primary) prevention; targeted (secondary) prevention, and tertiary prevention, focusing primarily on preventing mild mental health problems and promoting well-being in the workplace rather than treating mental disorders. It was specifically designed for SMEs and public agencies (for a comprehensive view of the design and rationale of the current study, please see references [,]), filling a critical gap identified in recent literature. EMPOWER integrates organizational readiness assessments and user-centered design principles to enhance feasibility and effectiveness, aiming to improve outcomes such as presenteeism, anxiety, depression, and well-being. We conducted a cluster randomized controlled trial (RCT) with employees from SMEs and public agencies in Finland, Poland, Spain, and the United Kingdom to evaluate the effectiveness of EMPOWER on various mental health and work-related outcomes, both immediately postintervention and at follow-up, and explored several factors that could impact the effectiveness through subgroup analyses. The rationale behind using a cluster randomized design was the organizational nature of the intervention and the high risk of contamination between employees within the same workplace. The primary and secondary outcomes were assessed at the individual employee level, while randomization and intervention allocation were performed at the organizational (cluster) level. Additionally, we discuss the challenges encountered during the implementation process and interpret our results in light of these difficulties.
This study was conducted as a multicountry cluster RCT with 2 parallel arms (intervention and wait-list control), following an initial stepped-wedge trial design in Spain, Finland, Poland, and the United Kingdom. Clusters (ie, companies or departments) were randomly assigned to A, B, and C groups, differing in the time when participants had access to the intervention. Participants from the 3 groups were invited to download the application at the same time, that is, at the beginning of the RCT (T0). Participants in group A had access to the intervention after completing the baseline questionnaires (T0), participants in group B had access after completing 2 rounds of questionnaires (after T1), and participants in group C had access after completing 3 rounds of questionnaires (after T2). A total of 2 additional rounds of questionnaires were delivered to the participants as follow-up measures (T3 and T4). As such, data were collected at the individual level at 5 time points []. Each data collection wave took place with a span of approximately 7 weeks of difference. The first study protocols were approved by the ethics committees of Fundació Sant Joan de Déu (PIC-39‐20), Turku University Hospital (PIC-993966082), University of York (HSRGC250321), and Institute of Occupational Medicine, University of Lodz (9/2020).
Periodic reports and meetings were held among the consortium partners to monitor the progress of the fieldwork. We faced several barriers and challenges throughout the trial period, particularly during the first months (February 2022 to September 2022). During this time, various contingency measures were implemented to increase participation, including the inclusion of new companies, improvements in local dissemination activities, and sending reminders to workers. In Table S1 in , we present the contextual barriers (at the country, company, and individual levels) as well as the facilitators identified throughout the trial.
In September 2022, the number of recruited participants was still lower than expected, despite the contingency actions that we took (see section Sample Size Calculation; 461 had consented to participate, 295 had completed baseline questionnaires, 100 had answered T1, 45 had answered T2, and 14 had answered T3). Feedback from participants, collected through qualitative interviews, revealed that one of the main reasons for not completing the planned assessments was the length of the questionnaires and the high number of assessment points. After careful consideration, the consortium agreed to modify the design from a stepped-wedge to a cluster RCT with 2 parallel arms: an intervention arm and a wait-list control arm. Following this change, each newly recruited cluster (company or department) was randomized to either intervention or control, regardless of size, type of workers, or sector. We also reduced the number of assessment points from 5 to 3 in all groups: T0 (baseline), T1 (postintervention, 7 weeks), and T2 (follow-up, 21 weeks). In addition, the number of questionnaires at T1 was reduced to minimize participant burden.
Participants recruited before September 2022, and therefore under the initial stepped-wedge design, were reallocated to either the intervention or control group depending on their study status. Specifically, participants from group A, who had already completed the intervention between T0 and T1, were directly assigned to the intervention group. Participants from group B were reassigned to the intervention group, with their T1 data treated as preintervention, T2 as postintervention, and their follow-up postponed by 7 weeks to ensure comparability with group A. Finally, participants in group C continued their original assessment schedule, with the exception that they did not gain access to the intervention until the follow-up assessment (21 weeks).
This study’s design modification was approved by Fundació Sant Joan de Déu in October 2022. The trial was registered at ClinicalTrial.gov (NCT04907604). The trial ended as planned after completion of follow-up assessments.
This cluster RCT was reported in accordance with the CONSORT (Consolidated Standards of Reporting Trials) 2010 statement [], the CONSORT extension for cluster randomized trials [], and the CONSORT extension for abstracts [] ().
Recruitment Process and RandomizationBefore the recruitment of participants, we approached several SMEs and public agencies from Finland, Poland, Spain, and the United Kingdom to participate in this study. Large companies were also considered for the RCT. Clusters were eligible if they were SMEs, public agencies, or large organizations willing to participate in this study and able to provide access to their employees for the duration of the trial. There was no limitation on the size or economic sector. Participant recruitment took place between February 25, 2022, and September 30, 2023; data collection lasted until May 30, 2024. New companies or departments could be included during the course of this study.
Different strategies were used by the local research teams to recruit companies (such as press release and media, employers’ organizations, and face-to-face and online meetings). A formal agreement was signed between the EMPOWER consortium and the employer before recruiting their employees. In a first round of randomization (January 2022) and under the stepped-wedge trial design, clusters (ie, SMEs or departments) were randomly allocated into group A, B, or C, considering strata to reduce imbalance issues []. This was done separately in each country, considering, when possible, the cluster size, type of workers (blue vs white collar), and type of company (SME, public agency, and large). A second round of randomization was repeated in April 2022. We used the runiform module from Stata (StataCorp LLC) to randomize newly recruited companies or departments. After the design modification with 2 parallel groups (from September 2022), newly recruited clusters were randomly allocated to the intervention or control group by using randomly generated numbers in Excel (Microsoft Corp). In this case, variables such as cluster size, type of workers, and type of company were not taken into account when randomizing the companies. Randomization was performed by this study’s coordinator. Moreover, due to the nature of the digital intervention, neither participants nor researchers were blinded to group allocation. Allocation concealment was not feasible due to the cluster-level nature of the intervention and the pragmatic implementation of the trial.
A total of 25 clusters were assigned to the intervention group and 19 to the control group. Initially, email invitations to the employees using the EMPOWER email address were sent, which provided a nontransferable, one-use link for downloading the application and a company-specific code to access it. This allowed us to have control over participants, but at the same time, did not allow participants to share the link with other workmates and increased the chance of technical problems. Thus, in April 2022, we created QR codes for each country to download the application, together with a company code to sign in. All procedures followed the General Data Protection Regulation [].
Participant inclusion criteria were (1) being aged 18 years or older, (2) having a mobile phone with internet access, (3) having sufficient knowledge of the local language, and (4) providing informed consent.
All participants were invited to download the application and provide informed consent, which was integrated into the application. After providing consent, participants were asked to answer the baseline assessment questions. Employees of localities assigned to the intervention group were able to access and use the intervention part of the application directly after completing the baseline questionnaires; application access lasted for a duration of 7 weeks. All participants were asked to complete additional rounds of questionnaires 7 weeks after baseline assessment completion (postintervention) and 21 weeks after postintervention assessment completion (follow-up). Once the 3 rounds of questionnaires were completed, participants in both groups (intervention and control groups) had full access to the intervention part of the application. Participants with severe mental health conditions (including people reporting suicidal thoughts) requiring immediate treatment were advised to seek professional help and were warned that the EMPOWER application was not an intervention; however, they were not excluded from this study.
The assessment protocols were delivered through the application with no time limit for participants to complete the questionnaires. Participants received automatic reminder emails and/or pop-up messages from the application to notify them when an assessment protocol was available for them to complete. To increase retention rates, weekly reports were used to follow up on the number of responses, and the local researchers actively sent personalized email reminders to participants to encourage them to answer the questionnaires. Local researchers in each country had a coordination handbook that covered topics such as study design, data management and security, ethical aspects, responsibilities and good practices, solving problems during the fieldwork, and the incidental findings policy. No specific procedures were implemented to monitor adverse events, as the intervention was low-intensity and self-guided. Any potential unintended effects were assessed descriptively. Participants were provided with the contact details of the local research teams.
Sample Size CalculationA cautionary approach was adopted to detect small changes and effect sizes and was based on a review that assessed the impact of mental health programs on presenteeism in the workplace []. The calculation also accounted for potential loss to early and late follow-ups (25%). The sample size calculation was based on detecting changes in presenteeism, based on a review of the impact of mental health programs on presenteeism in the workplace []. Considering an expected effect size of d=0.30, a bilateral type I error of 0.05, and a power of 80%, a total sample of 729 participants was required. However, due to the implementation of multilevel analyses resulting from cluster randomization, the sample size needed to be adjusted by a factor of 1.2, leading to a final sample size of 874 participants [].
Ethical ConsiderationsThis study was conducted in compliance with ethical guidelines for research involving human participants. Ethical approval was obtained from the relevant ethics committees in each participating country: Clinical Research Ethics Committee of Fundació Sant Joan de Déu (PIC-39-20), Ethics Committee of the Hospital District of Southwest Finland (PIC-993966082), Health Sciences Research Governance Committee, University of York (HSRGC250321), and the Institute of Occupational Medicine and the Bioethics Committee of the Nofer Institute of Occupational Medicine (Łódź, Poland; 9/2020). The trial was registered at ClinicalTrials.gov (NCT04907604). This study’s design modifications made in response to recruitment challenges were approved by Fundació Sant Joan de Déu in October 2022.
All participants provided informed consent before enrollment in this study. Consent was integrated into the EMPOWER application, where participants were required to review study information and explicitly agree to participate before accessing the baseline questionnaire. Participants had the right to withdraw from this study at any time without consequence. Participants who completed all 3 rounds of questionnaires (baseline, post intervention at 7 wk, and follow-up at 21 wk) received a €20 (US $22.94) voucher as compensation for their time and effort in this study.
All data collected in this study were anonymized and stored securely in compliance with the General Data Protection Regulation. Each participant was assigned a unique identification code to ensure confidentiality. Identifiable personal data was not collected, and all responses were encrypted during data transmission and storage. Employers had access only to aggregated, anonymized workplace risk assessments, ensuring that individual participant data remained confidential.
InterventionThe EMPOWER digital intervention has been previously described in detail []. The intervention included components delivered at both the individual employee level (self-guided application modules) and the organizational level (employer web portal and aggregated psychosocial risk feedback). The website is public and offers a company-level mental health awareness campaign that assists both employees and employers in addressing work-related stress and mental health challenges. The application is available for use by individual employees and contains triage assessment tools and an algorithm to tailor content modules based on their symptomatology and work functioning. The application follows a self-guided approach, and the content is based on cognitive behavioral therapy (CBT) techniques, including psychoeducational material, tracking your mood, self-guided goal setting, problem solving, relaxation, and breathing exercises. The different modules of the application were designed to increase awareness about mental health, reduce stress, anxiety, and other psychological symptoms, screen for psychosocial risk factors and recommendations, and include a return-to-work module. From an organizational level, employers could access a restricted web portal that provided an anonymous brief assessment of psychosocial risk factors in the workplace as reported by employees, accompanied by tailored recommendations for employers. The readiness of the website and application was evaluated using the technology readiness level adapted for implementation sciences [] as part of the impact assessment, which is described elsewhere []. During the RCT, we collected feedback from employees, employers, and local researchers as part of a qualitative assessment and impact analysis, described elsewhere [].
Outcome MeasuresAll primary and secondary outcomes were assessed at the individual employee level. A summary of the outcome measures used in this study can be found in . Levels of anxiety and depression were categorized using cutoff scores from the GAD-7 (Generalized Anxiety Disorder Questionnaire) and PHQ-9 (Patient Health Questionnaire-9) scales, where scores greater than 5 are considered as having at least mild anxiety symptoms [] or mild depressive symptoms [].
Table 1. Variables and measures for this study.OutcomeMeasureDescriptionPrimary outcome measurePresenteeismiMTA Productivity Cost Questionnaire (iPCQ) []Scores start at 0 and measure total hours lost due to being unproductive at workSecondary outcome measuresDepressionPatient Health Questionnaire-9 (PHQ-9) []Its total score ranges from 0 to 27, ranging from no depression to severe depression.AnxietyGeneralized Anxiety Disorder Questionnaire (GAD-7) []This questionnaire ranges from 0 to 21, meaning from no anxiety to severe anxiety.InsomniaInsomnia Severity Index (ISI) []The ISI ranges from 0 to 12, where 0 indicates no presence of insomnia, and 12 indicates a severe presence of insomnia.Work stressMini-Psychosocial Stressors at Work Scale (Mini-PSWS) []The total score ranges from 0 (no stress) to 48 (maximum stress)General stressPerceived Stress Scale (PSS-4) []The total score of the PSS-4 ranges from 0 to 16, where 0 indicates no stress and 16 indicates severe stress.Well-beingWorld Health Organization Well-Being Index (WHO-5) []It scores from 0 to 100, where higher scores indicate better well-being, and a clinically significant change was defined as a difference of 10 points or more.Momentaneous well-beingVisual analog scale from 0 to 10Higher scores represent better well-beingMental Health Quality of LifeMental Health Quality of Life Questionnaire (MHQoL) []Ranging from 0 to 21, where 21 represents the best possible mental health and quality of life.Physical activityInternational Physical Activity Questionnaire (IPAQ) []Reported as total METs, with higher values indicating greater activity levels.SittingInternational Physical Activity Questionnaire (IPAQ) []Was recorded in minutes per day spent sittingSomatizationPatient Health Questionnaire (PHQ-15) []Scores from 0 to 30, where 30 indicates the worst healthAbsenteeismiMTA Productivity Cost Questionnaire (iPCQ) []Scores start at 0 and measure total hours lost due to being absent from workUnpaid laboriMTA Productivity Cost Questionnaire (iPCQ) []It calculates the hours lost due to needing assistance with tasksaiMTA: Institute for Medical Technology Assessment.
bMET: metabolic equivalent of task.
Level of Engagement (Active vs Inactive Application Users)In our analysis of the intervention group, participants were categorized as either active or inactive application users based on several key criteria. For each participant, we evaluated whether they earned more than 5 badges, engaged with any content by reading pages, began working on any problems, consistently tracked their mood for more than a week (7 d), or acquired any badge through application interaction. Participants were assigned a score of 1 if a condition was met and 0 if it was not. We then summed these scores for each participant. Those who met none or only one of these conditions were classified as inactive, while participants who met more than 1 condition were classified as active.
Statistical AnalysisFor descriptive analysis, we used mean and SD, and frequencies and percentages (%) for continuous and categorical variables, respectively. Differences between groups were calculated using chi-squared and t tests. We also calculated a correlation matrix between the main outcomes at baseline to assess the relationships between those variables (see Table S2 in ).
As a first step, we followed an intention-to-treat approach to maintain the comparability between participants in the intervention and the control groups and avoid selection bias []. We calculated linear mixed models (LMMs) to separately analyze the primary and secondary outcomes. The independent variables were the participant’s group (intervention vs control) and the intervention time (baseline, postintervention, and follow-up). Each model was adjusted for the covariates gender, age, and country to account for potential confounding factors. Additional analyses were also performed, adding the same baseline measure as a covariate, to control for potential group differences at baseline. For each outcome, a P value and CIs are provided. LMMs were calculated for the whole sample. To investigate potential moderators for the effectiveness of the intervention on several outcomes, we conducted subgroup analyses where LMMs were estimated separately for different subgroups: gender, age groups (18‐35, 36‐49, and 50+ years), low and high levels of anxiety, depressive and somatization symptoms (using the clinical cutoff point of 5), low and high levels of work stress (using median as cutoff). For each subgroup, LMMs were adjusted for sex, age, and country to account for potential confounding factors (these covariates were held constant in each of the models; information about the coefficients for each covariate in the LMM models is available upon request).
Statistics that compared participants who dropped out with those who did not were also described with mean (SD) or n (%), depending on whether the outcomes were numeric or categorical. Finally, we conducted a per-protocol analysis as a secondary analysis, in which we conducted LMMs differentiating between active and inactive participants (see section Level of Engagement (Active vs Inactive Application Users)), comparing them to the control condition and between each other. As such, the per-protocol analyses in our study did not handle missing data. Per-protocol analysis estimates the effect of receiving the treatment (efficacy). Missing data were handled by using LMMs, which allow inclusion of participants with incomplete data under the missing at random assumption.
A total of 877 employees accepted the informed consent, and 708 answered the baseline questionnaires. A total of 44 clusters were randomized, with 25 allocated to the intervention group and 19 to the control group. No clusters were lost or excluded after randomization, and all of them were included in the primary analyses. A total of 333 participants completed postintervention questionnaires and 267 follow-up questionnaires (). shows the descriptive statistics for the total sample and by intervention group, 78% (544/697) were females, and the average age was 45.05 (SD 10.50) years, 18% (127/708) of participants were from Spain, 7% (51/708) from Poland, 55% (389/708) from the United Kingdom, and 20% (141/708) from Finland. Most participants were working in public agencies (584/685, 85%) and were white-collar workers (663/696, 95%). The intervention and control groups significantly differed for sex, country, company, and workers’ type. All participants who completed baseline assessments were included in the intention-to-treat analyses for the primary outcome.
Figure 1. Participant flow diagram. Table 2. Sociodemographic sample descriptive statistics for the total sample and by intervention group.VariableTotal sampleControl groupIntervention groupDifferencesTest (df)P valueGender, n (%)8.33 (2).02Female544 (78.05)292 (82)252 (73.47)Age, mean (SD)44.20 (10.52)44.17 (10.03)44.23 (11.01)0.126 (688.8).88Country, n (%)22.538 (3).001Spain127 (17.94)48 (13.30)79 (22.77)Poland51 (7.20)17 (4.71)34 (9.80)United Kingdom389 (54.94)225 (62.33)164 (47.26)Finland141 (20)71 (20)70 (20)Company type, n (%)39.017 (2).001Public agency584 (85.26)327 (90.58)257 (79.32)SME90 (13.14)23 (6.37)67 (20.68)Large company11 (1.61)11 (3.05)0 (0)Blue or white collar, n (%)13.338 (1).001White collar663 (95.26)347 (98.30)316 (92.13)aChi-square test.
bt test.
cSME: small to medium enterprise.
In general, the correlation matrix showed high correlation coefficients between most outcomes at baseline (except for the variables sitting and absenteeism, see Table S2 in ). shows the mean values for the outcomes at preintervention, generally, and divided by intervention and control groups. At preintervention (T0), the intervention group had significantly lower levels of depressive symptoms (t=2.031, P=.043), anxiety symptoms (t=2.169, P=.03), and general stress (t=3.104, P=.002), and higher levels of well-being (t=−3.412, P=.001), mental health quality of life (t=−2.166, P=.031), and momentaneous well-being (t=−2.438, P=.015), as compared with the control group. No statistically significant differences between groups were found for insomnia, work stress, physical activity, sitting, somatic health, absenteeism, presenteeism, and unpaid labor (all P>.05).
Table 3. Baseline means (SDs) and group comparisons for primary and secondary outcome variables.VariableT0TotalCGIGt test value (degrees of freedom)P valueDepressive symptoms, mean (SD)7.26 (5.72)7.71 (5.59)6.81 (5.82)2.031 (671.8).04Anxiety symptoms, mean (SD)6.10 (4.96)6.52 (4.84)5.68 (5.06)2.169 (653.9).03Insomnia symptoms, mean (SD)3.52 (2.60)3.63 (2.61)3.41 (2.59)1.106 (645.3).27Work stress, mean (SD)10.42 (9.75)10.37 (9.54)10.46 (9.94)−0.112 (616.9).91General stress, mean (SD)6.00 (3.24)6.39 (3.11)5.61 (3.32)3.104 (648.6).002Well-being, mean (SD)49.44 (21.47)46.54 (21.41)52.27 (21.29)−3.412 (640.1).001MH Quality of Life, mean (SD)13.88 (3.32)13.59 (3.37)14.15 (3.25)−2.166 (634.6).03Momentaneous well-being, mean (SD)6.56 (1.97)6.36 (2.03)6.74 (1.90)−2.438 (631.4).015Physical activity, mean (SD)6903.26 (6313.60)6461.78 (5998.93)7388.73 (6619.26)−1.776 (566.9).08Sitting, mean (SD)405.13 (248.54)407.85 (241.78)402.30 (255.70)0.294 (689.6).77Somatic complaints, mean (SD)7.66 (4.32)8.00 (4.32)7.32 (4.30)1.989 (632.4).047Absenteeism, mean (SD)717.88 (1803.00)641.90 (1684.23)791.01 (1910.14)−1.040 (623.6).30Presenteeism, mean (SD)8.39 (19.41)10.00 (22.83)6.84 (15.28)2.037 (535.1).04Unpaid labor, mean (SD)8.98 (33.40)10.58 (38.66)7.44 (27.37)1.174 (553).24aCG: control group.
bIG: intervention group.
Adjusted Main Model for Postintervention EffectsAll analyses were conducted following an intention-to-treat approach, including all participants according to their original group allocation. The results from models for the different outcomes adjusted by age, gender, and country are displayed in . The displayed value is the regression coefficient for the interaction effect from group (control vs intervention) by time (pre- to postintervention), therefore representing the intervention effect. Overall, the intervention had no significant effect on any of the outcomes of this study (all P>.05). Moreover, due to the significant differences at baseline between the intervention and control group, we also conducted the analyses adjusting for baseline levels. These additional analyses led to similar results except for a lower well-being in the intervention group as compared to the control group (β=−3.289, P=.034, 95% CI −6.331 to −0.247); however, this reduction in well-being is not considered clinically relevant.
Table 4. Generalized linear model adjusted results for postintervention effects.Outcomeβ-value (95% CI)P valueDepressive symptoms−0.052 (−1.010 to 0.905).92Anxiety symptoms−0.328 (−1.168 to 0.512).44Insomnia symptoms−0.196 (−0.633 to 0.240).38General stress0.385 (−0.195 to 0.965).19Work stress0.526 (−1.656 to 2.708).64Well-being−3.274 (−6.974 to 0.425).42MH Quality of Life−0.221 (−1.016 to 0.574).59Momentaneous well-being0.013 (−0.443 to 0.470).95Physical activity−224.285 (−2346.582 to 1898.012).84Sitting−50.343 (−124.856 to 25.169).19Somatic complaints0.289 (−0.389 to 0.947).39Absenteeism−55.672 (−439.783 to 328.439).78Presenteeism2.186 (−2.424 to 6.796).35Unpaid labor−0.04 (−8.356 to 3.653).99aAdjusted by sex, age, and country.
bMH: mental health.
Adjusted Main Model for Follow-Up EffectsAs there was no limit to when the participants could respond to the follow-up questionnaire, the follow-up time had a mean value of 48.84 (SD 24.77) weeks. When looking at the intervention effect at follow-up, no significant intervention effects were found for any of the outcome variables (all P>.05; see ). Again, due to the significant differences at baseline between the intervention and control group, we conducted the analyses adjusting for baseline (values or levels), finding no significant differences between the intervention and control group at follow-up.
Table 5. Generalized linear model adjusted results for follow-up effects.Outcomeβ-value (95% CI)P valueDepressive symptoms0.202 (−0.840 to 1.245).70Anxiety symptoms0.375 (−0.537 to 1.287).42Insomnia symptoms−0.148 (−0.619 to 0.323).54General stress0.123 (−0.502 to 0.749).70Work stress−0.610 (−2.387 to 1.166).50Well-being−1.022 (−5.019 to 2.975).62MH Quality of Life−0.211 (−0.945 to 0.524).57Momentaneous well-being−0.162 (−0.583 to 0.259).45Physical activity404.012 (−1219.670 to 2127.693).65Sitting−34.104 (−96.639 to 28.432).28Somatic health0.535 (−0.179 to 1.249).14Absenteeism−76.124 (−489.616 to 337.368).72Presenteeism1.294 (−3.608 to 6.396).58Unpaid labor3.653 (−5.322 to 12.629).42aAdjusted by sex, age, and country.
bMH: mental health.
Results From Subgroup AnalysesAnalyses conducted separately by subgroups yielded several significant results (tables with all the outcome variables for the different subgroup analyses can be found in Tables S3-S8 in ).
First, comparing men and women, we only found that the intervention group had decreased well-being at postintervention (β=−4.796, P=.029, 95% CI −9.109 to −0.482) in women. No significant intervention effects were found in men. The subgroup analysis by age groups showed that participants aged 18‐35 years were less sedentary, spending less time sitting at postintervention (β=283.325, P=.004, 95% CI −477.588 to −89.062). Participants aged 50 years and older showed increased presenteeism at postintervention (β=10.016, P=.007, 95% CI 2.714 to 17.317). Comparing participants with low and high depression scores at baseline showed that participants with low depressive levels had decreased well-being (β=−4.901, P=.029, 95% CI −9.304 to −0.497) and lower mental health quality of life (β=−0.964, P=.028, 95% CI −1.822 to −0.105) at postintervention. Conversely, participants with high depressive symptoms at baseline spent less time sitting at T1 (β=−147.023, P=.047, 95% CI −292.164 to −1.882). As for low and high levels of anxiety, we observed that those with low anxiety levels at baseline also experienced decreased well-being (β=−4.523, P=.048, 95% CI −9.011,−0.036) and reduced mental health quality of life (β=−1.246, P=.007, 95% CI −2.148 to −0.344) at T1. Participants with high anxiety levels had increased somatization at follow-up (β=1.413, P=.025, 95% CI 0.181 to 2.646). As for low and high levels of somatization at baseline, participants with low somatization at baseline were found to have a decrease in mental health quality of life at T1 (β=−2.095, P=.003, 95% CI −3.498 to −0.692), whereas they experienced less insomnia symptoms at follow-up (β=−0.855, P=.026, 95% CI −1.606 to −0.103). Participants with high levels of somatization at baseline had decreased well-being (β=−4.489, P=.049, 95% CI −8.962 to −0.017) and spent less time sitting at T1 (β=−103.532, P=.029, 95% CI −196.746 to −10.318). Finally, considering low and high levels of work stress, individuals with low work stress were found to report less well-being at follow-up (β=−6.673, P=.042, 95% CI −13.102 to −0.244), while individuals with high work stress reported less insomnia at follow-up (β=−0.822, P=.024, 95% CI −1.534 to −0.110).
Attrition AnalysisThe attrition analysis revealed several significant baseline differences between the participants who dropped out from this study and those who did not drop out (Table S9 in ). Age was slightly lower in the dropout group (mean 44.17, SD 10.03) compared to the not-dropout group (mean 44.23, SD 11.01), with a significant difference (t=3.201, P=.001). There were also significant differences in country distribution (χ²=13.135, P=.004), with a higher proportion of participants from the United Kingdom and a lower proportion from Spain and Poland in the dropout group. Depressive symptoms, anxiety symptoms, general stress, and work stress levels were significantly higher in the dropout group (depressive symptoms: t=−2.092, P=.037; anxiety symptoms: t=−2.784, P=.005; general stress: t=−3.101, P=.002; work stress: t=−2.133, P=.033). Absenteeism was also higher in the dropout group (t=−2.942, P=.003). Other variables did not show significant differences between the groups. Moreover, of the participants in the intervention group, only 34% were considered active users of the intervention, implying an overall low intervention adherence.
Per-Protocol AnalysesDue to high dropout rates and low engagement, we performed additional analyses to assess the impact of those biases. We divided participants in the intervention group into active and inactive users depending on their level of interaction with the application. We conducted LMM analyses and found that, again, none of the outcome measures was significantly different from the control group for both inactive and active groups, and for postintervention and follow-up measurement time (all P>.05, Table S10 in ). Additionally, when comparing active participants to inactive participants, no significant differences in any of the outcomes were found (all P>.05, Table S11 in ).
Sample Size and Sensitivity AnalysisThe final sample comprised 708 participants (347 in the intervention and 361 in the control group). Considering the corrective factor of 1.2 used in the a priori calculation (to account for the cluster design), the effective sample sizes are approximately 289 and 303 per group. With these numbers, the minimum detectable effect size for 80% power and 2-sided α=0.05 is 0.24-0.25 []. Thus, the current study would show adequate power to detect small-to-moderate effects (≥0.25-0.30) but would be underpowered for very small effects (≤0.20).
In a per-protocol comparison between intervention users (n=119) and controls (n=363), the effective intervention sample would drop to approximately 99 participants after the corrective factor, yielding a minimum detectable effect size of d between 0.34-0.35. This would considerably reduce the statistical power to detect small effects and should therefore be taken into account when interpreting our null findings.
Harms or Unintended EffectsNo adverse events or unintended effects related to the intervention were reported.
In this cluster RCT, we investigated whether a multimodal, digital platform, consisting of a website and an application, improved mental health, well-being, and work-related outcomes of employees from SMEs and public agencies in 4 European countries.
Effectiveness ResultsOverall, we did not find evidence of effectiveness of the intervention in the short (7 wk) or long term (21 wk post intervention), but some positive effects were observed for certain outcomes and subgroups of employees. We did not observe the expected effectiveness of the intervention to improve depressive symptoms, anxiety symptoms, general and work-related stress, insomnia symptoms, physical activity, well-being, presenteeism, or absenteeism. However, different effect patterns were found when analyzing different subgroups. Previous literature showed small, significant effects of digital interventions in the workplace []. In the Carolan et al [] meta-analysis (2017), 9 of the 21 studies included did not find an effect of the digital intervention on mental health. Additionally, 8 of 13 included studies did not find an intervention effect on work outcomes. It is possible that longer follow-up assessments (more than 21 weeks) are needed to detect an effect of an intervention like that proposed by EMPOWER.
As previously stated, when analyzing the effectiveness of the intervention in different subgroups, we found some significant results. For younger participants (18‐35 y), those with high levels of depressive symptoms at baseline and with high baseline somatization, the intervention showed a positive effect to reduce sedentary behavior. Additionally, insomnia improved at follow-up for participants with low baseline somatization. However, in other subgroups, the intervention had the opposite effect. These findings suggest that demographic factors and initial health conditions could significantly influence the effectiveness and outcomes of interventions, highlighting the need for tailored approaches in mental health and well-being programs. Previous evidence has shown that tailored digital interventions did not reduce anxiety and depressive symptoms in the general working population, but significantly alleviated these conditions in employees with higher psychological distress []. More research is needed to determine which characteristics of employees could moderate the effectiveness of OeMH interventions.
Contrary to our expectations, some subgroups (eg, women, participants with low baseline depressive and anxiety symptoms, and those with high baseline somatization) showed a short-term decrease in well-being post intervention, which disappeared at follow-up. Temporary declines in well-being and mental health after a psychological intervention are not uncommon and could result from increased self-awareness of one’s own issues and challenges. This is in line with the prevalence inflation hypothesis, which suggests that greater introspection and symptom monitoring can increase perceived distress before improvement occurs [].
Indeed, the prevalence inflation hypothesis posits that heightened awareness during interventions can lead to increased reporting of mild symptoms as significant issues, potentially exaggerating perceived prevalence. According to this, while mental health awareness efforts are beneficial as they lead to more accurate reporting of previously unrecognized symptoms, these efforts may also cause individuals to interpret and report mild distress as mental health problems, potentially exacerbating symptoms through self-fulfilling behaviors and creating a cyclical and intensifying effect []. One could speculate that immediate benefits of an intervention such as EMPOWER might not always be observable and thus, longer follow-up assessments are needed to observe positive effects.
A recent review [] suggests that personalized digital interventions demonstrated positive outcomes when it comes to reducing presenteeism, alleviating stress, improving sleep, and addressing physical symptoms associated with somatization. Nevertheless, their effectiveness in reducing absenteeism is somewhat limited. Moreover, another systematic review reported that the content of the intervention based on stress management and mindfulness seems more effective than CBT-based content. This suggests that CBT may be less useful for preventive approaches and for workplace settings []. All in all, the findings suggest that a universal approach to digital mental health interventions may not be effective in addressing the nuanced needs of workplace populations. Future implementations should consider a targeted approach, focusing on individuals with expressed mental health challenges or higher baseline distress.
Challenges During the RCTGiven the various challenges we faced during the trial, we cannot dismiss the possibility that the EMPOWER results may be attributable to chance rather than reflecting the effectiveness of the EMPOWER intervention. Some of these challenges included external contextual factors, such as the COVID-19 pandemic or the Ukraine war, organizational barriers, complexity of the initial design and length of questionnaires, and technical problems with the application affecting accessibility and adherence.
On the one hand, several contextual challenges were encountered during the RCT. For instance, in Finland, there were large reforms of health care, social welfare, and rescue services at the beginning of 2023. In Poland, there was a lack of company involvement in promoting the project among employees and a fear of discussing mental health with employees. In the United Kingdom, there were several strikes, and in Spain, we observed a certain level of stigma in the workplace to openly talk about mental health in the workplace. Overall, the COVID-19 pandemic hindered the successful completion of this trial, and the Ukraine war has had a significant impact on the small and medium companies both in Finland and in Poland.
On the other hand, we faced several challenges related to this study’s design and the technical problems of the application itself. During the design phase, we conducted a prepilot study with volunteers who tested the beta version of the application []. However, there was no proper validation of the prototype in a real-world context (ie, employees), which hindered the early identification of
Comments (0)