judgement ai jugement algorithme

Beyond the algorithm: preserving human judgement in humanitarian crises

Annesha Mahanta
Annesha MahantaAnnesha Mahanta is a humanitarian analyst with over six years of experience supporting crisis analysis, humanitarian coordination and information management across Sudan, Ukraine, Lebanon, Somalia or the Occupied Palestinian Territories. She has supported organisations including the United Nations Office for the Coordination of Humanitarian Affairs, the United Nations Disaster Assessment and Coordination and the United Nations Development Coordination Office on humanitarian response and analysis initiatives in complex emergencies. She also leads training courses and capacity-building initiatives in humanitarian analysis and emerging technologies in the sector.
Madigan Johnson
Madigan JohnsonMadigan Johnson is a digital expert specialising in user experience, research, technology and communications. She has a Master’s degree in International Humanitarian Action from the Network on Humanitarian Action and worked in the private tech sector before returning to the humanitarian sector, bringing sharp product and program leadership experience that now shapes her work at the intersection of data, AI and humanitarian impact. At Data Friendly Space, Madigan spearheaded the organisation’s research on AI use in humanitarian contexts.

When an artificial intelligence ingests government propaganda as crisis data, or aggregates Gaza and the West Bank as a single operational reality, the error is more than technical. Drawing on live deployments in Myanmar and the Occupied Palestinian Territories, the two authors make the case for hybrid intelligence – not as a compromise, but as the only responsible path.


In March 2025, an earthquake struck Myanmar. Data Friendly Space’s GANNET SituationHub automatically ingested sources into its database. Among the ingested content was media later iden­tified as government propaganda. When the Myanmar Information Management Unit flagged this content during manual review, Data Friendly Space (DFS) staff overrode the AI’s data ingestion and ex­cluded the propaganda from analytical outputs that would inform humanitar­ian response decisions. This incident encapsulates a critical challenge: how do humanitarians harness the analyt­ical power of artificial intelligence (AI) while preserving the contextual intelli­gence and ethical judgement that define effective crisis response?

Large language models (LLMs) and retrieval-augmented generation (RAG) systems can process information at scales beyond human capacity, identify patterns in fragmented crisis data and accelerate the preparation of situation reports needed for aid allocation. Yet risks mirror promises. When trained on biased, incomplete or contextually inappropriate data, AI systems propa­gate falsehoods at scale, reinforce ste­reotypes and introduce false certainty into genuinely ambiguous situations. In settings where decisions determine resource allocation for affected popu­lations, uncritical reliance on AI carries high human stakes.

This article provides DFS’s account of im­plementing hybrid intelligence through our human-in-the-loop (HITL) approach, analysing failures and corrective ac­tions from SituationHub deployments in Myanmar and the Occupied Palestinian Territories. The central argument is that responsible AI in humanitarian settings requires neither blind algorithmic trust nor the complete rejection of insights, but rather deliberate, people-led prac­tices that leverage the complementary strengths of humans and technology.

GANNET’s architecture and control dynamics

GANNET SituationHub operates through a RAG system. Unlike autonomous web-crawling systems, GANNET’s infor­mation environment is bounded: every source enters through human curation, creating initial quality control. Yet curation does not eliminate bias, error or contex­tual misapplication. Continuous human oversight remains essential to ensure al­gorithmic outputs align with humanitarian principles and operational realities.

In HITL systems, AI drives inference and decision-making; humans intervene to provide corrections and supervision. GANNET SituationHub enables an HITL approach in which human analysts retain authority over analytical outputs, while GANNET supports information aggrega­tion and synthesis.[1]Tathagata Chakraborti, Sarath Sreedharan and Subbarao Kambhampati, “The Emerging Landscape of Explainable Automated Planning & Decision Making”, 29th International joint conference on … Continue reading This distinction clari­fies accountability. When humans remain ultimate decision-makers, responsibility for harm rests with human judgement rather than algorithmic operation.

However, HITL systems can present distinct vulnerabilities. Human inter­pretation of algorithmic output intro­duces biases that differ from those present in the data. Humans can misun­derstand confidence scores, over-rely on algorithmic summaries and accept AI-generated content as plausible, de­spite hallucinations.[2]In the field of AI, we call a hallucination a false or misleading response presented as a certain fact, for example a bibliographic reference [editor’s note]. See Geetha Aradhyula, … Continue reading Effective hybrid intelligence requires not only the re­tention of human authority but also a comprehensive understanding of both AI capabilities and limitations. Without this understanding, humans risk be­coming passive validators rather than active decision-makers. This distinc­tion between structural authority and substantive competence determines whether hybrid intelligence operates as intended or reduces human oversight to a mere formality.

Case analysis: detecting and correcting failure

Following the March 2025 earthquake, the Myanmar GANNET SituationHub automatically ingested sources to syn­thesise a situation overview. The sys­tem’s ability to rapidly aggregate and cross-reference sources provides clear value by bringing out data that informs resource-allocation decisions. Yet one critical limitation is that such systems can inherit biases and errors from source documents.

Myanmar: when propaganda mimics legitimacy

Government propaganda, when present­ed with linguistic and formatting pat­terns similar to those of legitimate crisis information, evades initial detection. GANNET lacked mechanisms to distin­guish credible media from manipulated government communications masquer­ading as crisis updates. The propaganda entered analytical output before human detection caught it. Had this content shaped humanitarian decision-making unchecked, it would have eroded com­munity trust and spread false informa­tion about crisis realities. The reactive nature of the override, detected after ingestion rather than prevented before­hand, reveals a fundamental vulnera­bility: content-filtering must account for sophisticated disinformation that mimics legitimate sources, not merely for technical formats.

This revealed a key weakness in RAG sys­tems: the reliance on the trustworthiness of its sources. In environments where in­formation ecosystems have been com­promised, this assumption becomes hazardous and demonstrates that source curation alone is insufficient. While the SituationHub limits its information en­vironment, it cannot anticipate which sources might be compromised or how the system tags that information. For ex­ample, the system cited a United Nations Children’s Fund (UNICEF) Humanitarian Action for Children (HAC) document, but pulled the information from a similar re­port for Cameroon due to source inges­tion and tagging errors at the document level. When such errors occur, analysts escalate to the technical team for adjust­ment. Identifying nuanced and reliable sources is therefore critical in contexts such as Myanmar, where information is often shaped by political agendas, propaganda and competing narratives.

Here, the HITL element is particularly valuable, as analysts with contextual knowledge of Myanmar can better as­sess source credibility, recognise polit­ically charged or misleading language and identify information that more ac­curately reflects realities on the ground.[3]Ana Beduschi, “Harnessing the potential of artificial intelligence for humanitarian action: opportunities and risks”, International review of the Red Cross, vol. 104, no. 919, June 2022, p. … Continue reading

Occupied Palestinian Territories: language as ethical boundary

In early deployments of the SituationHub in Occupied Palestinian Territories (OPT), analytical output employed language that was technically accurate but in­appropriate for humanitarian analysis. Media coverage in GANNET’s training sources emphasised regional language conventions, rather than terminology and framings appropriate to the OPT humanitarian analysis.

Different organisations use different lexicons to describe the same events. What one source calls “occupation forces”, another calls “Israeli military”. What one frames as “deliberate target­ing”, another frames as “military oper­ations”. Equally, sources aligned with other parties may employ terms such as “resistance fighters”, “martyrs” or “lib­eration operations”, framing carrying equally charged political valency in the opposite direction. These asymmetries are not incidental; they reflect the con­tested political landscape. A concrete illustration of this is the deliberate pol­icy of not naming specific organisations, an explicit decision to engage with hu­manitarians operating in the context to protect staff security, preserve human­itarian access and maintain relations with de facto authorities. However, that carries its own epistemic cost.

These instances raise a fundamental question about the ethics of intention­al constraint. Decisions to restrict a system’s vocabulary, whether by exclud­ing an organisation’s name or by avoid­ing a term that objectively describes a party’s actions, are not neutral technical choices. They are actions, made in full knowledge that descriptive accuracy is being sacrificed to preserve humanitar­ian access, protect personnel and main­tain relationships with authorities. Yet it introduces a tension that HITL systems in political contexts cannot resolve al­gorithmically: the tension between the humanitarian duty to do no harm and the moral imperative to report conditions truthfully. Human judgement must bear that tension consciously, not outsource it to a system prompt.

A related failure concerned geographic granularity. Operating across the oc­cupied Palestinian territory, GANNET consistently aggregated information from Gaza and the West Bank into uni­fied analytical output, treating them as a single operational reality. In a context where Gaza faces acute emergency con­ditions requiring immediate response and the West Bank faces challenges of a structurally different character, this conflation was not a minor imprecision but a harm-generating error: it obscured the severity of conditions in Gaza, mis­represented the nature of needs in the West Bank and denied decision-makers accurate, differentiated information. The correction required a structural interven­tion: a reconfigured system to operate at the governorate level. This case demon­strates that geographic resolution is not merely a technical parameter; it is an ethical one. The unit of aggregation de­termines whose needs become visible.

Corrections in these instances involved two mechanisms: immediate manual edit­ing before publication and medium-term technical collaboration to refine system prompting and source selection. Isolated errors were addressed on an ad hoc ba­sis, while systematic patterns were doc­umented and escalated by analysts and given to the technical team for long-term adjustment. This process highlights a crit­ical insight: hybrid intelligence requires continuous learning cycles in which hu­mans identify systematic failures, collab­orate with technologists to determine root causes and implement improve­ments to prevent recurrence.

Systemic vulnerabilities: why these failures matter

Both cases illustrate broader risks in humanitarian AI: algorithmic confidence masks crucial data gaps and contextual uncertainties.[4]Mirca Madianou, “Nonhuman humanitarianism: when ‘AI for good’ can be harmful”, Information, Communication & Society, vol. 24, no. 6, 2021, p. 850-868, … Continue reading AI systems are trained to produce fluent, coherent output that reads as authoritative. This is not a tech­nical flaw amenable to engineering solu­tions; it is a fundamental feature of LLMs.

In humanitarian contexts, conflating al­gorithmic fluency with factual accuracy carries particular danger. Humanitarians often have to make decisions based on incomplete information and educated judgement. When AI systems produce synthesised output that sounds fluent and authoritative, they risk leading decision-makers to treat them as more reliable than warranted.

“In humanitarian contexts, conflating algorithmic fluency with factual accuracy carries particular danger.”

Both cases illustrate that algorithms trained on heterogeneous data produce statistically defensible but ethically inappropriate output. Language bias in AI systems has been extensively docu­mented across medical, legal and hir­ing contexts.[5]Tino Kreutzer, James Orbinski, Lora Appel et al., “Ethical implications related to processing of personal data and artificial intelligence in humanitarian crises: a scoping review”, BMC Medical … Continue reading In humanitarian settings, language bias carries particular weight because humanitarian princi­ples are themselves linguistic com­mitments. Neutrality, impartiality and independence are commitments to par­ticular framings and vocabularies.[6]Andrea Guillén and Emma Teodoro, “Embedding Ethical Principles into AI Predictive Tools for Migration Management in Humanitarian Action”, Social Sciences, vol. 12, no. 2-53, 2023, … Continue reading

Geographic bias in humanitarian AI sys­tems presents a significant challenge.[7]Caroline M. Gevaert, Thomas Buunk and Marc J. C. van den Homberg, “Auditing Geospatial Datasets for Biases: Using Global Building Datasets for Disaster Risk Management”, IEEE Journal of Selected … Continue reading Systems trained predominantly on data from the Global North often struggle to perform effectively in Global South contexts where data is less abundant and local knowledge is less digitised. Information that is appropriate in one context may be unsuitable in anoth­er. These situations reflect structural inequalities regarding which informa­tion is digitised and whose knowledge is considered reliable. These inequalities are inherent to AI systems and require deliberate mitigation strategies.

Implementing hybrid intelligence: operational practices

GANNET demonstrates that hybrid intel­ligence works, but requires deliberate design. Three critical practices emerged from implementation.

Cross-functional reviews

Before output is published, analysts examine recommendations and review with the technical team. This is not mere quality control but collaborative learn­ing, often conducted in partnership with local and national responders. Analysts bring contextual expertise and usage knowledge; technologists understand system behaviour and constraints.

Mutual respect for different areas of expertise forms the foundation; neither technologists nor analysts solve these problems alone.

Tiered decision architecture

Not all output carries equal stakes. Low-stakes background information might appropriately rely more heavily on algorithmic synthesis, with human re­view limited to flagging obvious errors. High-stakes output, that which directly informs resource allocation or policy, re­quires deeper engagement. GANNET im­plementations can route output through different oversight levels based on de­cision impact, matching human effort to the magnitude of the consequences.

Trust-building

Algorithms excel at structured digital information; they struggle with tacit knowledge, relationships and informal understanding developed through com­munities and long-term engagement. Hybrid systems should incorporate human intelligence, observations, rela­tionships and cultural sensitivities into system prompts and decision frame­works. Systematising what was previ­ously intuitive creates accountability and repeatability.

Hybrid intelligence requires organisa­tional commitment to treating human override as a learning opportunity rath­er than as a system failure. DFS treated these occurrences as essential data for system behaviour rather than dismissing them as anomalies. The team systemat­ically documented patterns, collabo­rated with technologists to determine causes and implemented improvements. Organisations that discourage dissent from algorithmic recommendations will inevitably receive less accurate hu­man oversight.[8]Jerry Alan Fails and Dan R. Olsen, Jr., “Interactive Machine Learning”, 8th International conference on intelligent user interfaces, Miami, Fl., 12-15 January 2003, p. 39-45, … Continue reading

Grounding hybrid intelligence in humanitarian principles

Effective hybrid intelligence cannot be purely technical. It must be grounded in humanitarian principles that serve as de­cision criteria when human judgement conflicts with an algorithmic recommen­dation. The humanitarian principles of humanity, impartiality, neutrality and independence cannot be automated because they require contextual judge­ment informed by relationships and ethical reasoning. Humanity requires an understanding of what constitutes harm in specific cultural and politi­cal contexts; no algorithm can sub­stitute for human moral judgement. Impartiality requires understanding whether particular language choices, data sources or framing systematically disadvantage populations; algorithms lack inherent commitments to fair treatment.

“No algorithm can substitute for human moral judgement.”

Neutrality requires knowing which framings and sources align with established political positions; this knowledge is fundamentally contextual and relational. Independence requires recognising subtle state interference, including sophisticated disinformation; algorithms cannot assess credibility in political contexts.[9]International Committee of the Red Cross, The fundamental principles of the Red Cross: commentary, analysis from Jean Pictet, 1979, … Continue reading

When properly implemented, hybrid intelligence strengthens humanitarian principles by processing information at scale, freeing analysts to focus on judgement-intensive tasks. However, these benefits depend on humans main­taining genuine control, understanding system strengths and limitations, identi­fying and correcting failures, and retain­ing the authority to override AI outputs when humanitarian principles demand it. The true measure of hybrid intelli­gence in humanitarian analysis is not simply improved analysis but analysis aligned with humanitarian values.

Neither tool nor oracle

Responsible AI in humanitarian crises re­quires humans as active decision-makers, not passive algorithmic validators. The cases demonstrate that human oversight is non-negotiable and requires deliber­ate effort. Oversight emerges from or­ganisational practices building in review points, cross-functional collaboration and failure-based learning.

Polarised debates on AI, positioning it as a solution or a threat, often obscure a more pragmatic reality: AI is a tool with specific capabilities and limitations. Systems, organisations and practices must ensure that humans remain gen­uinely in control. This means more than retaining decision authority; it means humans understand the systems they deploy, identify failures and estab­lish mechanisms to learn from errors. Practices must align with humanitar­ian principles, ensuring that systems support, rather than undermine, our commitment to them. AI should be re­garded as one tool among many that serve people affected by crises, rather than a replacement for the contextual intelligence, ethical reasoning and re­lational knowledge that define human­itarian action.

The path forward is structured HITL sys­tems grounded in humanitarian princi­ples, designed to augment, rather than replace, judgement. This approach is neither simple nor foolproof. It requires investment in co-design, training, organi­sational culture and technical infrastruc­ture. It requires sustained commitment to treating failures as learning opportu­nities rather than system defects. But the alternative, wholesale rejection of AI’s potential or deployment of algorithmic systems without meaningful oversight serves neither humanitarian principles nor crisis-affected people. The choice is not whether to use AI, but how to pre­serve human judgement.

 

Picture credit: Sandra Calligaro / Picturetank / ACF

Support Humanitarian Alternatives

Was this article useful and did you like it? Support our publication!

All of the publications on this site are freely accessible because our work is made possible in large part by the generosity of a group of financial partners. However, any additional support from our readers is greatly appreciated! It should enable us to further innovate, deepen the review’s content, expand its outreach, and provide the entire humanitarian sector with a bilingual international publication that addresses major humanitarian issues from an independent and quality-conscious standpoint. You can support our work by subscribing to the printed review, purchasing single issues or making a donation. We hope to see you on our online store! To support us with other actions and keep our research and debate community in great shape, click here!

References

References
1 Tathagata Chakraborti, Sarath Sreedharan and Subbarao Kambhampati, “The Emerging Landscape of Explainable Automated Planning & Decision Making”, 29th International joint conference on artificial intelligence [proceedings], Yokohama, Japan, January 2021, p. 4803-4811, https://www.ijcai.org/Proceedings/2020/669
2 In the field of AI, we call a hallucination a false or misleading response presented as a certain fact, for example a bibliographic reference [editor’s note]. See Geetha Aradhyula, “Human-in-the-Loop Learning Systems”, World Journal of Advanced Engineering Technology and Sciences, vol. 11, no. 1, 2024, p. 514-521.
3 Ana Beduschi, “Harnessing the potential of artificial intelligence for humanitarian action: opportunities and risks”, International review of the Red Cross, vol. 104, no. 919, June 2022, p. 1149-1169, https://international-review.icrc.org/sites/default/files/reviews-pdf/2022-06/harnessing-the-potential-of-artificial-intelligence-for-humanitarian-action-919.pdf
4 Mirca Madianou, “Nonhuman humanitarianism: when ‘AI for good’ can be harmful”, Information, Communication & Society, vol. 24, no. 6, 2021, p. 850-868, https://www.tandfonline.com/doi/full/10.1080/1369118X.2021.1909100
5 Tino Kreutzer, James Orbinski, Lora Appel et al., “Ethical implications related to processing of personal data and artificial intelligence in humanitarian crises: a scoping review”, BMC Medical Ethics, vol 26, no. 49, 2025, https://doi.org/10.1186/s12910-025-01189-2
6 Andrea Guillén and Emma Teodoro, “Embedding Ethical Principles into AI Predictive Tools for Migration Management in Humanitarian Action”, Social Sciences, vol. 12, no. 2-53, 2023, https://www.mdpi.com/2076-0760/12/2/53
7 Caroline M. Gevaert, Thomas Buunk and Marc J. C. van den Homberg, “Auditing Geospatial Datasets for Biases: Using Global Building Datasets for Disaster Risk Management”, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, 2024, p. 12579-12590, https://ieeexplore.ieee.org/document/10584113
8 Jerry Alan Fails and Dan R. Olsen, Jr., “Interactive Machine Learning”, 8th International conference on intelligent user interfaces, Miami, Fl., 12-15 January 2003, p. 39-45, https://dl.acm.org/doi/10.1145/604045.604056
9 International Committee of the Red Cross, The fundamental principles of the Red Cross: commentary, analysis from Jean Pictet, 1979, https://www.icrc.org/en/article/fundamental-principles-red-cross-commentary

You cannot copy content of this page