Data for famine and extreme emergencies: challenges and opportunities

From a multisectoral stakeholder point of view, there are many competing interests in the analysis of famine or extreme emergencies that border on famine, that result in major challenges. The major actors in this narrative include national governments, donors, humanitarian agencies (both international and local), information systems, and affected populations. Each has its own perspective but may have very different power to control data collection and analysis. In general, there are several challenges common to all three types of analysis such as classification, causal path determination, and early warning forecasting. These common issues are briefly described below.

Data collection, availability, and accessibility

Procedures, protocols, and quality are among the major constraints in data collection. These limitations can leave entire categories of data missing, creating critical obstacles to sound analysis. The most frequently “missing” data categories include mortality and nutrition from Standardized Monitoring and Assessment of Relief and Transitions (SMART) surveys, and specifically, the prevalence estimates of malnutrition that are critical to meeting the mandated requirements for methodological rigor. Information on food security tends to dominate these analyses, but while famine involves food insecurity, it results from much more than just inadequate dietary intake data. Information from health, water, sanitation and hygiene (WASH) programs is also frequently missing.

But beyond these indicators that give a picture of current conditions and nutrition status, deeper causal analysis requires data on the drivers of the crisis (typically conflict, climate hazards and market hazards), as well as information on underlying vulnerability and mitigating factors. Monitoring climate and market conditions is relatively well established; monitoring conflict is less so, and as noted above, the presence of conflict makes all this difficult [5].

Historical data for the analysis of trends—an important component of contextual analysis—may be incomplete or non-existent or limited to specific subpopulations. The methodology of data collection, and therefore the collected data, may vary depending on the purpose for which data were collected. Famine analysis usually requires data that are representative of a general population, yet much humanitarian data are collected for program management purposes, and thus not necessarily representative of a population—only of the specific program participants.

Collecting information about mortality, malnutrition, vulnerability and coping can constitute an invasion of the privacy of very vulnerable populations. When the affected populations may live in areas under the control of armed groups fighting against the state, the compromising of such data may put people at physical risk. Ensuring the privacy and security of sensitive data is an equally important consideration, especially when dealing with vulnerable populations.

Access to affected populations is a major constraint, and a major reason for the data gaps noted above. Conflicts are the common causal factors across almost all contemporary famines or near-famine emergencies. Conflicts limit access to affected populations and thus to essential information needed for assessing current conditions (and thus for classification analysis) and understanding dynamics to enable forecasting. Access to affected populations may be constrained by actual insecurity, but it may also be constrained by bureaucratic requirements—particularly if the governments involved do not necessarily want current conditions to be publicized—whether there is conflict happening or another driver of the crisis. This can result in significant gaps in the data, making analysis very difficult.

There are five major knowledge and data gaps to build a robust data infrastructure and foster information technology advancements. As shown in Fig. 1, we recognize that 1— insufficient data granularity, 2—inadequate data integration, 3—limited community engagement, 4—structural missingness, and 5—limited data literacy are interconnected and mutually amplifying challenges.

Fig. 1Fig. 1

Five major knowledge and data gaps in humanitarian settings

Data sharing and timeliness

The ability for analysts to access all data on a timely basis is a major constraint. Data may exist, but unless it can be fed into an analysis—and that analysis can be cross-checked by other analysts—the data might not be very useful for detecting and preventing famine. Data sharing—or more specifically the reluctance to share data—was among the most frequently mentioned constraints to famine analysis more generally. Sometimes this may have to do with political considerations, but frequently it is more a worry that—given some of the constraints on data collection—data quality might not be great and thus reflect badly on the agency or team collecting the data. And finally, information is power, and whoever controls data tends to “control the narrative” on a given crisis or affected population. All of these are constraints to data sharing [4].

Data sharing is becoming more common, with platforms such as the Humanitarian Data Exchange (HDX) serving as a platform where data can be shared and aggregated. Yet, too often, data do not turn up in these locations until after its immediate usage for humanitarian planning and response (or early warning) has long since expired. Particularly regarding needs assessment or early warning, data are only useful for planning and response for a relatively limited time. Data might be useful for research purposes thereafter, but not for informing a timely response. As part of the Humanitarian Reset, triggered by the closure of USAID and other funding cuts, efforts are underway to make the timely sharing of data a more common practice—the impact has yet to be seen.

Data quality and manipulation of data and analysis

Data quality may suffer, particularly when access and time constraints are acute, and it may also be deliberately undermined or manipulated. IPC analysis in particular, but generally all forms of famine analysis, has quite high standards for data quality—many of which are very difficult to attain in conflict or access-constrained contexts. While data quality may be affected by incomplete or have lots of missing answers, poor sampling techniques, access and time constraints, limitations in enumerator training or supervision and a variety of other considerations. But sometimes there are also subtle and not so subtle attempts to subvert or manipulate data and data collection [6].

The reliance of data collection and analysis on funding sources is often raised as a constraint. The very agencies collecting and analyzing information are dependent on the results of the analysis for their budgets, creating at least the appearance of a conflict of interest. In many ways, IPC analysis—in which there are multiple sources of information and multiple parties to the analysis—is a check against this. Separating agencies that are tasked to collect data and conduct analysis (such as ACAPS, REACH-Impact Initiatives, or FEWS NET, the Famine Early Warning System Network—a US-funded famine information project) from implementing agencies is one way of addressing this issue. However, implementing agencies still dominate in data collection related to food security.

Much of the data used in famine classification analysis is collected and analyzed, at best, on a semi-annual basis—in many contexts only once per year, sometimes not at the most critical time of the year (which could be during the lean or pre-harvest season), and sometimes not even at the same time of year, making it difficult to compare across seasons, locations, livelihood zones, etc. Even with semi-annual assessments, these are large-scale efforts that involve multiple teams and lots of logistics—they are not something that can be conducted with sufficient frequency. Furthermore, a lot can change between these large-scale assessments, calling for more real-time or near-real-time monitoring and analysis [7].

Analytical constraints

Constraints to analysis may follow many of the same issues and they may be more overt as well. Examples of famine analysis being quashed or stopped by governments are rare (as famine itself is rare) yet have been documented. Perhaps more troubling is self-censorship by analysts to prevent intervention from governments. Most likely, the interest here is not just the degree of severity (as in IPC classifications) but the numbers related to ‘population in need’ or the so-called PIN numbers. This applies to current status assessment and forecasting future numbers, with the latter being the basis on which budgeting is to some extent based. Higher PIN numbers may mean more resources, but not so much so that they put numbers of affected people into Phase 5 or famine. One form of this was noted in the previously mentioned study, which showed increasing numbers of affected populations in each of the IPC phase classifications from Phase 1 through Phase 4, but none in Phase 5—the so-called ‘left-skewed but truncated’ distribution of population-in-need [4].

The use of qualitative information and expert judgment

Finally, and especially regarding understanding famine or extreme crisis dynamics, much of the information needed cannot be boiled down to quantitative data—even though current status assessment relies almost exclusively on quantitative data that vary in accuracy. Interpreting trends in terms of contextual background, such as livelihood systems, seasonality, national policy (and politics!), donor relations, and more, is all very important to understanding the dynamics that can lead to famine and extreme crisis. This in turn at least partially relies on expert judgment that is not easily translated into an algorithm [8]. Furthermore, expert judgment is often applied broadly, but the difficulty lies in making such judgments at a more granular level without local expertise. Conversely, expert judgment data are often not scalable unless local expertise is involved.

In preparing annual reports and assessments, there is a lack of agreement on the data aggregation methods across time and geospatial domains and across different regions and organizations. The issue of aggregating continuous data to categories whose boundaries may make sense but still retain an element of arbitrariness is not unique to famine and is quite dangerous in many fields. This may be exacerbated when continuous or discrete data with drastically different precision, accuracy, and coverage are mysteriously recalibrated and presented as a ‘magical’ score. This runs the risk of masking the complexity of underlying issues—such that what is a continuum of extreme outcomes is almost always reduced to a question of ‘if it or isn’t it’ (a famine, in this case). That is less a data question than an interpretation question, and it is compounded by dubious or non-existent data. Several recent studies have attempted to bridge the gap in famine modeling by relying more heavily on qualitative information and expert judgment [9, 10].

Misuse, misinformation, and disinformation

In the current stage of advancements in AI, misinformation and data manipulation are—and likely to continue to be—rapidly increasing and there are few current reliable safeguards and regulations. Data are political capital, with outright misinformation and disinformation, and with interested parties attempting to spin data and analysis (or what passes for ‘data and analysis’) to suit their agenda. There are some proposals on how governments or other entities can curb misinformation [11]. The need to combat misinformation requires redirecting already scarce resources. This often happens at the expense of building the infrastructure for data collection in affected areas and training skilled personnel to collect, process, and analyze data accurately.

Comments (0)

No login
gif