When an artificial intelligence ingests government propaganda as crisis data, or aggregates Gaza and the West Bank as a single operational reality, the error is more than technical. Drawing on live deployments in Myanmar and the Occupied Palestinian Territories, the two authors make the case for hybrid intelligence – not as a compromise, but as the only responsible path.
In March 2025, an earthquake struck Myanmar. Data Friendly Space’s GANNET SituationHub automatically ingested sources into its database. Among the ingested content was media later identified as government propaganda. When the Myanmar Information Management Unit flagged this content during manual review, Data Friendly Space (DFS) staff overrode the AI’s data ingestion and excluded the propaganda from analytical outputs that would inform humanitarian response decisions. This incident encapsulates a critical challenge: how do humanitarians harness the analytical power of artificial intelligence (AI) while preserving the contextual intelligence and ethical judgement that define effective crisis response?
Large language models (LLMs) and retrieval-augmented generation (RAG) systems can process information at scales beyond human capacity, identify patterns in fragmented crisis data and accelerate the preparation of situation reports needed for aid allocation. Yet risks mirror promises. When trained on biased, incomplete or contextually inappropriate data, AI systems propagate falsehoods at scale, reinforce stereotypes and introduce false certainty into genuinely ambiguous situations. In settings where decisions determine resource allocation for affected populations, uncritical reliance on AI carries high human stakes.
This article provides DFS’s account of implementing hybrid intelligence through our human-in-the-loop (HITL) approach, analysing failures and corrective actions from SituationHub deployments in Myanmar and the Occupied Palestinian Territories. The central argument is that responsible AI in humanitarian settings requires neither blind algorithmic trust nor the complete rejection of insights, but rather deliberate, people-led practices that leverage the complementary strengths of humans and technology.
GANNET’s architecture and control dynamics
GANNET SituationHub operates through a RAG system. Unlike autonomous web-crawling systems, GANNET’s information environment is bounded: every source enters through human curation, creating initial quality control. Yet curation does not eliminate bias, error or contextual misapplication. Continuous human oversight remains essential to ensure algorithmic outputs align with humanitarian principles and operational realities.
In HITL systems, AI drives inference and decision-making; humans intervene to provide corrections and supervision. GANNET SituationHub enables an HITL approach in which human analysts retain authority over analytical outputs, while GANNET supports information aggregation and synthesis.[1]Tathagata Chakraborti, Sarath Sreedharan and Subbarao Kambhampati, “The Emerging Landscape of Explainable Automated Planning & Decision Making”, 29th International joint conference on … Continue reading This distinction clarifies accountability. When humans remain ultimate decision-makers, responsibility for harm rests with human judgement rather than algorithmic operation.
However, HITL systems can present distinct vulnerabilities. Human interpretation of algorithmic output introduces biases that differ from those present in the data. Humans can misunderstand confidence scores, over-rely on algorithmic summaries and accept AI-generated content as plausible, despite hallucinations.[2]In the field of AI, we call a hallucination a false or misleading response presented as a certain fact, for example a bibliographic reference [editor’s note]. See Geetha Aradhyula, … Continue reading Effective hybrid intelligence requires not only the retention of human authority but also a comprehensive understanding of both AI capabilities and limitations. Without this understanding, humans risk becoming passive validators rather than active decision-makers. This distinction between structural authority and substantive competence determines whether hybrid intelligence operates as intended or reduces human oversight to a mere formality.
Case analysis: detecting and correcting failure
Following the March 2025 earthquake, the Myanmar GANNET SituationHub automatically ingested sources to synthesise a situation overview. The system’s ability to rapidly aggregate and cross-reference sources provides clear value by bringing out data that informs resource-allocation decisions. Yet one critical limitation is that such systems can inherit biases and errors from source documents.
Myanmar: when propaganda mimics legitimacy
Government propaganda, when presented with linguistic and formatting patterns similar to those of legitimate crisis information, evades initial detection. GANNET lacked mechanisms to distinguish credible media from manipulated government communications masquerading as crisis updates. The propaganda entered analytical output before human detection caught it. Had this content shaped humanitarian decision-making unchecked, it would have eroded community trust and spread false information about crisis realities. The reactive nature of the override, detected after ingestion rather than prevented beforehand, reveals a fundamental vulnerability: content-filtering must account for sophisticated disinformation that mimics legitimate sources, not merely for technical formats.
This revealed a key weakness in RAG systems: the reliance on the trustworthiness of its sources. In environments where information ecosystems have been compromised, this assumption becomes hazardous and demonstrates that source curation alone is insufficient. While the SituationHub limits its information environment, it cannot anticipate which sources might be compromised or how the system tags that information. For example, the system cited a United Nations Children’s Fund (UNICEF) Humanitarian Action for Children (HAC) document, but pulled the information from a similar report for Cameroon due to source ingestion and tagging errors at the document level. When such errors occur, analysts escalate to the technical team for adjustment. Identifying nuanced and reliable sources is therefore critical in contexts such as Myanmar, where information is often shaped by political agendas, propaganda and competing narratives.
Here, the HITL element is particularly valuable, as analysts with contextual knowledge of Myanmar can better assess source credibility, recognise politically charged or misleading language and identify information that more accurately reflects realities on the ground.[3]Ana Beduschi, “Harnessing the potential of artificial intelligence for humanitarian action: opportunities and risks”, International review of the Red Cross, vol. 104, no. 919, June 2022, p. … Continue reading
Occupied Palestinian Territories: language as ethical boundary
In early deployments of the SituationHub in Occupied Palestinian Territories (OPT), analytical output employed language that was technically accurate but inappropriate for humanitarian analysis. Media coverage in GANNET’s training sources emphasised regional language conventions, rather than terminology and framings appropriate to the OPT humanitarian analysis.
Different organisations use different lexicons to describe the same events. What one source calls “occupation forces”, another calls “Israeli military”. What one frames as “deliberate targeting”, another frames as “military operations”. Equally, sources aligned with other parties may employ terms such as “resistance fighters”, “martyrs” or “liberation operations”, framing carrying equally charged political valency in the opposite direction. These asymmetries are not incidental; they reflect the contested political landscape. A concrete illustration of this is the deliberate policy of not naming specific organisations, an explicit decision to engage with humanitarians operating in the context to protect staff security, preserve humanitarian access and maintain relations with de facto authorities. However, that carries its own epistemic cost.
These instances raise a fundamental question about the ethics of intentional constraint. Decisions to restrict a system’s vocabulary, whether by excluding an organisation’s name or by avoiding a term that objectively describes a party’s actions, are not neutral technical choices. They are actions, made in full knowledge that descriptive accuracy is being sacrificed to preserve humanitarian access, protect personnel and maintain relationships with authorities. Yet it introduces a tension that HITL systems in political contexts cannot resolve algorithmically: the tension between the humanitarian duty to do no harm and the moral imperative to report conditions truthfully. Human judgement must bear that tension consciously, not outsource it to a system prompt.
A related failure concerned geographic granularity. Operating across the occupied Palestinian territory, GANNET consistently aggregated information from Gaza and the West Bank into unified analytical output, treating them as a single operational reality. In a context where Gaza faces acute emergency conditions requiring immediate response and the West Bank faces challenges of a structurally different character, this conflation was not a minor imprecision but a harm-generating error: it obscured the severity of conditions in Gaza, misrepresented the nature of needs in the West Bank and denied decision-makers accurate, differentiated information. The correction required a structural intervention: a reconfigured system to operate at the governorate level. This case demonstrates that geographic resolution is not merely a technical parameter; it is an ethical one. The unit of aggregation determines whose needs become visible.
Corrections in these instances involved two mechanisms: immediate manual editing before publication and medium-term technical collaboration to refine system prompting and source selection. Isolated errors were addressed on an ad hoc basis, while systematic patterns were documented and escalated by analysts and given to the technical team for long-term adjustment. This process highlights a critical insight: hybrid intelligence requires continuous learning cycles in which humans identify systematic failures, collaborate with technologists to determine root causes and implement improvements to prevent recurrence.
Systemic vulnerabilities: why these failures matter
Both cases illustrate broader risks in humanitarian AI: algorithmic confidence masks crucial data gaps and contextual uncertainties.[4]Mirca Madianou, “Nonhuman humanitarianism: when ‘AI for good’ can be harmful”, Information, Communication & Society, vol. 24, no. 6, 2021, p. 850-868, … Continue reading AI systems are trained to produce fluent, coherent output that reads as authoritative. This is not a technical flaw amenable to engineering solutions; it is a fundamental feature of LLMs.
In humanitarian contexts, conflating algorithmic fluency with factual accuracy carries particular danger. Humanitarians often have to make decisions based on incomplete information and educated judgement. When AI systems produce synthesised output that sounds fluent and authoritative, they risk leading decision-makers to treat them as more reliable than warranted.
“In humanitarian contexts, conflating algorithmic fluency with factual accuracy carries particular danger.”
Both cases illustrate that algorithms trained on heterogeneous data produce statistically defensible but ethically inappropriate output. Language bias in AI systems has been extensively documented across medical, legal and hiring contexts.[5]Tino Kreutzer, James Orbinski, Lora Appel et al., “Ethical implications related to processing of personal data and artificial intelligence in humanitarian crises: a scoping review”, BMC Medical … Continue reading In humanitarian settings, language bias carries particular weight because humanitarian principles are themselves linguistic commitments. Neutrality, impartiality and independence are commitments to particular framings and vocabularies.[6]Andrea Guillén and Emma Teodoro, “Embedding Ethical Principles into AI Predictive Tools for Migration Management in Humanitarian Action”, Social Sciences, vol. 12, no. 2-53, 2023, … Continue reading
Geographic bias in humanitarian AI systems presents a significant challenge.[7]Caroline M. Gevaert, Thomas Buunk and Marc J. C. van den Homberg, “Auditing Geospatial Datasets for Biases: Using Global Building Datasets for Disaster Risk Management”, IEEE Journal of Selected … Continue reading Systems trained predominantly on data from the Global North often struggle to perform effectively in Global South contexts where data is less abundant and local knowledge is less digitised. Information that is appropriate in one context may be unsuitable in another. These situations reflect structural inequalities regarding which information is digitised and whose knowledge is considered reliable. These inequalities are inherent to AI systems and require deliberate mitigation strategies.
Implementing hybrid intelligence: operational practices
GANNET demonstrates that hybrid intelligence works, but requires deliberate design. Three critical practices emerged from implementation.
Cross-functional reviews
Before output is published, analysts examine recommendations and review with the technical team. This is not mere quality control but collaborative learning, often conducted in partnership with local and national responders. Analysts bring contextual expertise and usage knowledge; technologists understand system behaviour and constraints.
Mutual respect for different areas of expertise forms the foundation; neither technologists nor analysts solve these problems alone.
Tiered decision architecture
Not all output carries equal stakes. Low-stakes background information might appropriately rely more heavily on algorithmic synthesis, with human review limited to flagging obvious errors. High-stakes output, that which directly informs resource allocation or policy, requires deeper engagement. GANNET implementations can route output through different oversight levels based on decision impact, matching human effort to the magnitude of the consequences.
Trust-building
Algorithms excel at structured digital information; they struggle with tacit knowledge, relationships and informal understanding developed through communities and long-term engagement. Hybrid systems should incorporate human intelligence, observations, relationships and cultural sensitivities into system prompts and decision frameworks. Systematising what was previously intuitive creates accountability and repeatability.
Hybrid intelligence requires organisational commitment to treating human override as a learning opportunity rather than as a system failure. DFS treated these occurrences as essential data for system behaviour rather than dismissing them as anomalies. The team systematically documented patterns, collaborated with technologists to determine causes and implemented improvements. Organisations that discourage dissent from algorithmic recommendations will inevitably receive less accurate human oversight.[8]Jerry Alan Fails and Dan R. Olsen, Jr., “Interactive Machine Learning”, 8th International conference on intelligent user interfaces, Miami, Fl., 12-15 January 2003, p. 39-45, … Continue reading
Grounding hybrid intelligence in humanitarian principles
Effective hybrid intelligence cannot be purely technical. It must be grounded in humanitarian principles that serve as decision criteria when human judgement conflicts with an algorithmic recommendation. The humanitarian principles of humanity, impartiality, neutrality and independence cannot be automated because they require contextual judgement informed by relationships and ethical reasoning. Humanity requires an understanding of what constitutes harm in specific cultural and political contexts; no algorithm can substitute for human moral judgement. Impartiality requires understanding whether particular language choices, data sources or framing systematically disadvantage populations; algorithms lack inherent commitments to fair treatment.
“No algorithm can substitute for human moral judgement.”
Neutrality requires knowing which framings and sources align with established political positions; this knowledge is fundamentally contextual and relational. Independence requires recognising subtle state interference, including sophisticated disinformation; algorithms cannot assess credibility in political contexts.[9]International Committee of the Red Cross, The fundamental principles of the Red Cross: commentary, analysis from Jean Pictet, 1979, … Continue reading
When properly implemented, hybrid intelligence strengthens humanitarian principles by processing information at scale, freeing analysts to focus on judgement-intensive tasks. However, these benefits depend on humans maintaining genuine control, understanding system strengths and limitations, identifying and correcting failures, and retaining the authority to override AI outputs when humanitarian principles demand it. The true measure of hybrid intelligence in humanitarian analysis is not simply improved analysis but analysis aligned with humanitarian values.
Neither tool nor oracle
Responsible AI in humanitarian crises requires humans as active decision-makers, not passive algorithmic validators. The cases demonstrate that human oversight is non-negotiable and requires deliberate effort. Oversight emerges from organisational practices building in review points, cross-functional collaboration and failure-based learning.
Polarised debates on AI, positioning it as a solution or a threat, often obscure a more pragmatic reality: AI is a tool with specific capabilities and limitations. Systems, organisations and practices must ensure that humans remain genuinely in control. This means more than retaining decision authority; it means humans understand the systems they deploy, identify failures and establish mechanisms to learn from errors. Practices must align with humanitarian principles, ensuring that systems support, rather than undermine, our commitment to them. AI should be regarded as one tool among many that serve people affected by crises, rather than a replacement for the contextual intelligence, ethical reasoning and relational knowledge that define humanitarian action.
The path forward is structured HITL systems grounded in humanitarian principles, designed to augment, rather than replace, judgement. This approach is neither simple nor foolproof. It requires investment in co-design, training, organisational culture and technical infrastructure. It requires sustained commitment to treating failures as learning opportunities rather than system defects. But the alternative, wholesale rejection of AI’s potential or deployment of algorithmic systems without meaningful oversight serves neither humanitarian principles nor crisis-affected people. The choice is not whether to use AI, but how to preserve human judgement.
Picture credit: Sandra Calligaro / Picturetank / ACF

