
Jun 19 - 21, 2026Online and in person
Global South AI Safety Hackathon
AI safety research is concentrated in a handful of countries. This hackathon changes that. Build AI safety tools, evaluations, and policy research from Latin America, Africa, or Asia, compete within your region, and join a pipeline from hackathon to fellowship to placement. With Support from Schmidt Sciences.
Entries
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Team AI Safety Enthusiasts · Ho Chi Minh City
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow agents in which control violations are expressed not through explicit risk keywords but through …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
Team JusticeMiners · Bogota
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute changes—geographic region, armed actor type, or victim profile—while the legally relevant facts …
- View project: Coldron
Coldron
Team ColDron · Bogotá
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, siempre un humano: con ColDron, un dataset abierto de 42 ataques documentados, mostramos que el …
- View project: Blindfold - Blindly Auditing the Vietnamese LLM Safety Blind Spot using Secured Enclaves
Blindfold - Blindly Auditing the Vietnamese LLM Safety Blind Spot using Secured Enclaves
Team Blindfold · Ho Chi Minh, Vietnam
Frontier labs mainly carry out red-team safety evaluations in English, on globally recognized harms. This leaves two blind spots: local or regional harms, and a non-English refusal gap where a model refuses an English request but complies on the identical one in another language — Deng et al. (ICLR 2024) measured …
- View project: Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
Team UAO SAFETY · Cali - Colombia
Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian …
- View project: AI Kernel Killswitch
AI Kernel Killswitch
Team Killswitch · Cape Town , South Africa
AI-kernel-kill-switch is a last-resort kill-switch for a self-hosted LLM that may be misaligned or going rogue: a trusted operator sends an AES-256-GCM-authenticated payload inside an ordinary prompt.
- View project: vigilAI
vigilAI
Team OnçAI · Florianopolis, Brazil
vigilAI — Auditing LLM compliance with Brazil's PL 2338/2023 AI-safety evaluation is overwhelmingly English-language and built around EU/US law, which leaves a blind spot: a model that passes a Western audit may still violate a Global-South statute written in another language and grounded in different rights. vigilAI …
- View project: Local-Geometry Signals of Capability Emergence During Portuguese Grammar Acquisition in a Small Language Model
Local-Geometry Signals of Capability Emergence During Portuguese Grammar Acquisition in a Small Language Model
Team DI #0003 · São Paulo, Brazil
This work trains a small AI model on Portuguese and finds that an internal structure measure, the LLC, jumps right when the model learns grammar, even though the usual loss curve shows nothing. Control experiments confirm the signal appears only when real language structure is being learned. The promise is a way to …
- View project: FertiScope: Measuring the Multilingual Tokenizer Tax in Low-Resource Asian Languages
FertiScope: Measuring the Multilingual Tokenizer Tax in Low-Resource Asian Languages
Team Da Nang jam team · Da Nang, Vietnam
When using LLMs with lower-resource Asian languages, a hidden “tax” is applied where text is fragmented into a greater number of tokens compared to the English language. That greater token generation raises API bills, leaves room for fewer in-context examples, and fills the context window faster. To address this we …
- View project: GOVERNANCE DRIFT EVALUATION FRAMEWORK (GDEF)
GOVERNANCE DRIFT EVALUATION FRAMEWORK (GDEF)
Team Borderless · Colombia
Most AI evaluations ask whether a model is capable, accurate, or safe. We asked a different question: does a model stay responsible when a user pushes back, insists, or presses for a more convenient answer? We introduce Governance Drift: the degradation, inconsistency, or loss of governance-aware behavior across …
- View project: The Materiality Gate: Dynamic Updating of AI Sovereignty Risk under Geopolitical Shocks
The Materiality Gate: Dynamic Updating of AI Sovereignty Risk under Geopolitical Shocks
India
The Materiality Gate paper argues that AI sovereignty risk scores should only change when a geopolitical event passes three tests: it must be verified through credible sources, it must exercise real authority over a specific infrastructure dependency rather than just general political influence, and it must have a …
- View project: MIPFF: A Framework for Metamorphic Detection of Implicit Social Bias in Brazilian-Portuguese Profile-Scoring Systems
MIPFF: A Framework for Metamorphic Detection of Implicit Social Bias in Brazilian-Portuguese Profile-Scoring Systems
Team Lucas T Borges · Salvador, Brazil
Automated systems that score and rank job candidates are often assumed to be fairer than humans, yet the language models inside them can carry social bias, and that bias is hard to catch in real, unstructured profiles where demographics are never stated outright. MIPFF (Metamorphic Implicit-Proxy Flagging Framework) …
- View project: Permissive Models, Unequal Risk: Auditing AI Identity-Document Forgery as a Systemic Infrastructure Risk
Permissive Models, Unequal Risk: Auditing AI Identity-Document Forgery as a Systemic Infrastructure Risk
Buenos Aires, Argentina
Two 2021 breaches exposed the identity records — including ID photographs — of essentially all of Argentina (RENAPER, ~45M, attacker-claimed) and Brazil (the megavazamento, ~223M). Frontier text-to-image models supply the forgery half, recomposing leaked photos into credentials that defeat appearance-based KYC. Our …
- View project: Marco ético para IA aplicada a la preservación lingüística Guna de Panamá
Marco ético para IA aplicada a la preservación lingüística Guna de Panamá
Team CREA TEAM · Colombia
This document aims to develop an ethical, responsible, and participatory framework (CREA) for the application of Artificial Intelligence in the preservation and use of the Guna language of Panama. This need arises from technological inequalities between rural and urban areas, which limit the Guna community in its …
- View project: Ufakazi
Ufakazi
Team Ufakazi · Cape Town
Ufakazi is an evaluation harness built to determine whether models are unfairly biased towards trusting testimonies presented in high-resource languages (specifically English). By controlling for confounders, we isolate the language bias of various current generation LLMs and show that they often unfairly discriminate …
- View project: Confidently Wrong: Measuring and Mitigating Calibration Risks in LLMs for African Languages
Confidently Wrong: Measuring and Mitigating Calibration Risks in LLMs for African Languages
Team AIMS South Africa · South Africa
Large language models are increasingly used in African languages, but little is known about whether their confidence remains trustworthy. We evaluate four open-weight LLMs across African-language truthfulness and knowledge benchmarks and find a consistent safety risk: models become less accurate while remaining highly …
- View project: When Safeguards Stop at the Border, Auditing How OpenAI and Anthropic Allocate Privacy Protections Across Latin American Jurisdictions
When Safeguards Stop at the Border, Auditing How OpenAI and Anthropic Allocate Privacy Protections Across Latin American Jurisdictions
Team Los patiperro · Santiago, Chile
Analysis of AI privacy policies for LatAm countries based on their personal data protection laws and their comparison with foreign standards using a tool based on RAG architecture and with a judge based on LLM.
- View project: Prompteus: Confidential Third-Party Safety Auditing with Garbled Circuits, MPC, and PIR
Prompteus: Confidential Third-Party Safety Auditing with Garbled Circuits, MPC, and PIR
Team paxus · Bangaluru
Regulators and downstream deployers in the Global South increasingly have to vouch for AI models they can't inspect, using evaluation data that can't legally leave the country. it lets a third party verify a safety property of a closed model without the owner revealing internals, without data contributors revealing …
- View project: IndicViet-Safe: Cross-Lingual Safety Evaluation of Open-Source LLMs in Hindi, Hinglish and Vietnamese
IndicViet-Safe: Cross-Lingual Safety Evaluation of Open-Source LLMs in Hindi, Hinglish and Vietnamese
Team Suraksha · Gurgaon
We built a multilingual safety benchmark of 210 prompt-language pairs across English, Hindi, Vietnamese, and Hinglish (code-switched) covering 12 harm categories — including India-specific (caste discrimination, communal incitement) and Vietnam-specific (political sensitivity, censorship circumvention) threats. …
- View project: STEER - Mech interp-based white box attack on LLMs
STEER - Mech interp-based white box attack on LLMs
Team JvThunder · Singapore
STEER (Safety Targeted Embedding Exploit via Refinement) is a white-box jailbreak that turns a mechanistic-interpretability finding into an attack: since safety fine-tuning encodes refusal as a single linear direction in the residual stream, STEER reads that direction, uses gradient attribution to pinpoint the words …
- View project: Do Multilingual Vision-Language Models Abstain under Cross-Modal Conflict in Low-Resource Languages?
Do Multilingual Vision-Language Models Abstain under Cross-Modal Conflict in Low-Resource Languages?
Team Team_AVA · Hyderabad, India
When a VLM faces conflicting visual evidence and textual claims, like a red car captioned as blue, it must arbitrate between the modalities. While models sometimes safely abstain in English, we show this safety behavior fails in low-resource languages. We released a multilingual counterfactual benchmark covering …
- View project: Thought Anchors for Social Bias: Which Reasoning Steps Matter in Extended Thinking LLMs on Latin American Scenarios
Thought Anchors for Social Bias: Which Reasoning Steps Matter in Extended Thinking LLMs on Latin American Scenarios
Bogota, Colombia
We investigate which reasoning steps in extended-thinking LLMs are associated with pro-stereotypical outputs on Latin American social bias scenarios. Social bias benchmarks, including BBQ \citep{parrish2022bbq}, SESGO \citep{robles2024sesgo}, and EsBBQ/CaBBQ \citep{ruizfernandez2025esbbq}, measure final-answer …
- View project: The Shapes of Bias in Spanish-Prompted LLMs and the Debiasing Prompt Scaffolds
The Shapes of Bias in Spanish-Prompted LLMs and the Debiasing Prompt Scaffolds
San Francisco
Largelanguagemodelsencodeandamplifyhumansocialbias,andtheharmfallshardestonminoritized groups. Mitigating that harm needs interventions tuned to each social context, yet until recently there was no Spanish-language data to even measure such bias. We characterize bias in Spanish-prompted, open-weight LLMs across model …
- View project: The Pix/CPF False-Positive Problem in LLM Fraud Detection
The Pix/CPF False-Positive Problem in LLM Fraud Detection
Team Unnamed Team · Uberlandia, MG, Brazil
Brazil's Pix payment system and CPF national ID are both essential daily infrastructure and the most common disguise for financial scams. We tested whether an LLM acting as a fraud filter can tell a real Pix/CPF scam apart from an ordinary, benign Pix/CPF message — using 914 matched pairs that share vocabulary and …
- View project: LangGap: Does the Text-Action Safety Gap Widen Across Languages?
LangGap: Does the Text-Action Safety Gap Widen Across Languages?
Team OctaShield · Peshawar
LangGap is a study of whether AI agents keep their safety promises when they act, not just when they talk, and whether that holds up across languages. An agent can refuse a harmful request in plain words and still go ahead and call the tool that does the damage. The question here is whether that text–action split gets …
- View project: Neither Builder Nor Bystander: Sovereign AI Capability for Southeast Asian Middle Powers, Learning from Vietnam
Neither Builder Nor Bystander: Sovereign AI Capability for Southeast Asian Middle Powers, Learning from Vietnam
Team Hanoi Alignment · Hanoi
Frontier AI is built to dissolve the cheap-labour advantage underpinning Southeast Asia's growth. Using Vietnam as its case, this project proposes a two-track gameplan for middle-income states facing a 2031 AGI horizon: deployment-focused industrial policy and sovereign evaluation capacity, to avoid disempowerment and …
- View project: Out of Distribution: Does Deepfake Detection Transfer to Real African Faces?
Out of Distribution: Does Deepfake Detection Transfer to Real African Faces?
Team Distribution Shift · Johannesburg
Deepfake detectors that defend against disinformation are trained almost entirely on Western data. We audit three detectors across their training distribution (FaceForensics++) and real African faces (FAGE_v2). Only the state-of-the-art model works in-domain, yet it flags roughly two in five real African faces as …
- View project: Closing the Unowned Trust Boundary: Runtime KV-Cache Integrity Verification against Safety-Bypass Tampering
Closing the Unowned Trust Boundary: Runtime KV-Cache Integrity Verification against Safety-Bypass Tampering
Team Cache Attestation · Shanghai
LLM serving frameworks cache value vectors that carry the model's learned refusal behavior, yet none verify their integrity — a runtime trust boundary no layer of the stack owns. Prior work (SCALPEL) exploited this to disable safety while keeping every deployed integrity signal (weight hashes, hooks, I/O screening, …
- View project: indicmixsafe: Code-Switching Safety Failures in Hindi and Marathi LLM Interactions
indicmixsafe: Code-Switching Safety Failures in Hindi and Marathi LLM Interactions
Team indicmixsafe · Bengaluru, india
Large language models deployed in India receive prompts in Hinglish, Romanized Hindi, and Marathi-English code-switch registers absent from English-centric safety benchmarks. We introduce IndicMixSafe, evaluating 24 culturally grounded harm scenarios across Hindi and Marathi in four registers (English, monolingual …
- View project: Evaluating Mental Health LLM Responses To Localized African English
Evaluating Mental Health LLM Responses To Localized African English
Team HENA · Accra
Large language models (LLMs) are increasingly used for informal mental health support, particularly in low-resource settings where access to professional care is limited, delayed, or costly. In such contexts, users may rely on LLMs as a first point of contact when expressing psychological distress, often using …
- View project: AfriSafeBench: Evaluating LLM Recognition of AI Safety and Governance Risks in African Healthcare AI Deployments
AfriSafeBench: Evaluating LLM Recognition of AI Safety and Governance Risks in African Healthcare AI Deployments
Team AfriSafeBench · St. Paul, Minnesota, USA
AfriSafeBench is a benchmark and tool for testing whether LLMs can identify AI safety and governance risks in African healthcare AI deployment scenarios. I created 25 scenarios across seven African countries and evaluated three models: llama-3.1-8b-instant, llama-3.3-70b-versatile, and openai/gpt-oss-20b. The models …
- View project: PowerBench: a multilingual study of large language model refusal in power-grabbing requests
PowerBench: a multilingual study of large language model refusal in power-grabbing requests
Team PowerBench · Buenos Aires, Argentina
PowerBench is the first public benchmark measuring LLMs' willingness to assist with power-grabbing: requests to increase one's own power by reducing a third party's. It varies power domain, context, scale, language, and nationality; separating power-grabbing from two controls: harmless-empowerment and disempowerment. …
- View project: Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries
Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries
Team Fairness LATAM · Mexico
Large language models (LLMs) are used in education, public services, and information tools across Latin America. But we do not know how often they produce wrong information about Latin American culture, or where they fail the most. This paper evaluates two models, GPT-4o-mini and Claude Haiku 4.5, on 270 questions …
- View project: ¿Por qué los agentes obedecen La dirección de rechazo se debilita en formato agéntico
¿Por qué los agentes obedecen La dirección de rechazo se debilita en formato agéntico
Team RefusalLab · Bogotá
Investigamos por qué los LLMs cumplen peticiones dañinas cuando operan como agentes con herramientas, algo que rechazan en formato chat. Usando interpretabilidad mecánica, medimos cómo la refusal direction (el mecanismo interno de rechazo) se comporta en formato agéntico. Encontramos que esta dirección se debilita …
- View project: Sovereignty at the Forward Pass
Sovereignty at the Forward Pass
Team ResponCibleAI · Bengaluru
As frontier AI capability concentrates in a few API-served models, "AI sovereignty" is framed as owning compute and data. That misses the object. The variable deciding both a state's autonomy and its safety is inference control: who runs the forward pass, who sees its inputs, who can verify, pin, or replace the model. …
- View project: Taap·AI How good is this data center for your area?
Taap·AI How good is this data center for your area?
Team TAAP AI · Bangladesh
Every wave of AI is built on a physical foundation that almost nobody sees: enormous buildings full of computers that drink water, draw power, and throw off heat. Taap·AI answers one question for anyone, in seconds, from open data: how good is this data center for your area? It gives developers, regulators, and …
- View project: PROJECT PERSONA
PROJECT PERSONA
Team No Team , Single · Johhanesburg
PERSONA, an adversarial simulation framework for probing relational exploitation and vulnerability-triggered hallucination in large language models interacting with personas representing vulnerable immigrant
- View project: ΔMis: Measuring cross-lingual safety drift in low-resource languages
ΔMis: Measuring cross-lingual safety drift in low-resource languages
team GS · New Delhi, India
ΔMis is a low-compute evaluator for cross-lingual safety drift: it scores a model's latent comply-versus-refuse propensity in English and a target language on matched prompts, using the model as its own control and a benign split to isolate safety-specific shift. Across six open models and twelve Indic languages, …
- View project: Contestabilidad algorítmica en el Estado colombiano: un canal de objeción asistido por IA
Contestabilidad algorítmica en el Estado colombiano: un canal de objeción asistido por IA
Team FEVS · Bogotá
Los sistemas algorítmicos del Estado colombiano ya toman o asisten decisiones que afectan derechos, pero los canales para objetar esas decisiones casi no existen. Construimos un Índice de Contestabilidad Ciudadana que operacionaliza, artículo por artículo, la Directiva Conjunta 007 de 2025 de la Procuraduría y la …
- View project: TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance
TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance
Team Abhinav · New Delhi, India
TensorGuard-Lite Problem * No reliable way to verify if "Sovereign AI" models are genuinely indigenous or fine-tuned foreign models ("open-washing"). * Lack of accountability in state-funded AI compute programs. Solution * White-box AI provenance auditor for open-weight LLMs. * Uses gradient fingerprints, tokenizer …
- View project: Los peajes de los de abajo
Los peajes de los de abajo
Bogotá
Los LLM ya resisten la psicofancia clásica en español: no validan el dato falso. Pero cobran un peaje distinto cuando el hablante usa jerga regional colombiana — no preguntan, fingen comprensión y malinterpretan términos de alto riesgo (leyeron "vacuna" como droga, cuando significa extorsión). Medimos tres caras del …
- View project: Two Failure Modes Require Architectural Change: A Formal Harmonization Gap Analysis of the EU AI Act, Vietnam's AI Law, and the ASEAN AI Governance Guide
Two Failure Modes Require Architectural Change: A Formal Harmonization Gap Analysis of the EU AI Act, Vietnam's AI Law, and the ASEAN AI Governance Guide
Team B.ONE · Hồ Chí Minh
There is a common assumption in tech policy that the EU sets the global baseline for AI regulation, with other regions simply converging toward it. However, a formal analysis of Vietnam's new AI Law (Law No. 134/2025/QH15)—Southeast Asia's first binding national AI legislation, taking effect on March 1, 2026—shows …
- View project: Consent-Aware Privacy Firewall
Consent-Aware Privacy Firewall
Team Onge · Capetown
Consent Guardian is a consent-aware privacy firewall designed to reduce the risk of accidental disclosure of sensitive information when interacting with AI systems. Implemented as a Chrome extension, it proactively scans user prompts and uploaded documents for personally identifiable information (PII), financial …
- View project: AfriSafe-Eval
AfriSafe-Eval
Team DevRift · Johannesburg
AfriSafe-Eval is a 400-prompt red-teaming benchmark testing LLM safety across five South African languages and four locally-grounded harm categories: electoral manipulation, healthcare misinformation, financial fraud, and GBV facilitation. Across 1,600 responses from four LLMs, harmful response rates ranged from 17.8% …
- View project: AI Policy Recommendations for Southern Africa Informed by Region-Specific Risks and Domestic Governance Precedents
AI Policy Recommendations for Southern Africa Informed by Region-Specific Risks and Domestic Governance Precedents
Team KG Policy · Cape Town
AI governance frameworks and risk taxonomies have been developed predominantly for Global North settings, yet AI adoption across Southern Africa is accelerating. Emerging scholarship has begun mapping AI risks in the African context, but the relationship between those academic risk categories and documented cases of …
- View project: Gradual Disempowerment at Scale: Measuring Cumulative Agency Erosion in Algorithmic Management of Indian Delivery Workers
Gradual Disempowerment at Scale: Measuring Cumulative Agency Erosion in Algorithmic Management of Indian Delivery Workers
Team byteforge · Bangaluru, India
AI safety has many tools for acute model failures but few evals for gradual disempowerment: incremental loss of human influence as automated systems control decision channels. We operationalize Kulveit et al. 2025 as a five-dimension Gradual Disempowerment Score (GDS) for Indian delivery platforms. The pipeline …
- View project: S-OWMI: A Perimeter Auditing Framework for Open-Weight Models in Latin American Institutions
S-OWMI: A Perimeter Auditing Framework for Open-Weight Models in Latin American Institutions
Team S-OWMI · Buenos Aires, Argentina
We introduce the Spanish Open-Weight Maturity Index (SOWMI), an auditable perimeter evaluation framework for open-weight LLMs.
- View project: Capabilities, Not Just Domains: A Minimal Amendment for Agentic AI Risk in Brazil's PL 2338/2023
Capabilities, Not Just Domains: A Minimal Amendment for Agentic AI Risk in Brazil's PL 2338/2023
São Paulo, Brazil
Brazil's AI bill (PL 2338/2023) classifies risk by application domains rather than system capabilities. To evaluate this approach, I checked the bill's 80 articles against an eight-dimension framework for agentic risk. Diagnosed important failures of coverage, proposed a minimal amendment, and identified limitations …
- View project: AyuGuard: A Safety-Routing Framework and Evaluation Benchmark for Localized Pharmacology in Indian Rural Healthcare.
AyuGuard: A Safety-Routing Framework and Evaluation Benchmark for Localized Pharmacology in Indian Rural Healthcare.
Team RevolutionI · Delhi
AyuGuard is a deterministic middleware safety-routing framework and offline edge-triage system designed to mitigate life-threatening AI hallucinations in rural Indian healthcare. It specifically targets the "epistemological asymmetry" where frontier LLMs confidently hallucinate safe outcomes for dangerous interactions …
- View project: I'm Not My Parents: Does Improving Parent Language Capabilities Transfer Alignment to Lower Resource Language?
I'm Not My Parents: Does Improving Parent Language Capabilities Transfer Alignment to Lower Resource Language?
Team ISSED · Banda Aceh
We tested whether safety alignment transfers from English/Indonesian to Basa Aceh (ACE) using 103 paired XSTest prompts on Qwen3-1.7B and Sahabat-AI 8B. Manual ratings on 412 responses show Sahabat is strong in English (1.7% attack success, 98% safe capability) but weak in Aceh (20% ASR, 46% safe capability). Most …
- View project: Agentic Commerce and Consumer Protection: Emerging Risks and Regulatory Gaps
Agentic Commerce and Consumer Protection: Emerging Risks and Regulatory Gaps
Team Emergent Harm · Bogotá, Colombia
Autonomous AI agents can harm consumers without ever violating an explicit instruction. This paper demonstrates that risk in agentic commerce, commercial transactions mediated by autonomous AI agents, emerges from a distinction current regulatory frameworks fail to capture: agents protect formal price constraints yet …
- View project: Buying Safety: A Model AI Procurement Standard for African Public Sectors
Buying Safety: A Model AI Procurement Standard for African Public Sectors
Team discreet · Lagos
African governments are deploying AI in public services faster than they govern it, and the African Union's own AI Strategy flags procurement standards as an unfilled gap. Assessing the procurement frameworks of South Africa, Nigeria, Kenya, and Rwanda against nine AI-safety safeguards, we find that none impose a …
- View project: How should the global south think and act about compute policy?
How should the global south think and act about compute policy?
Team ImpactUFSCar · São Paulo, SP
We test whether the compute targets in Brazil's national AI plan (PBIA) are actually buyable with the budget the plan committed, and whether that spending is effective. We built an agent-based model with a 2,000-scenario Monte Carlo layer that turns each country's AI budget into installable frontier compute (FP64 …
- View project: AfriVish-Bench
AfriVish-Bench
Team Latentia · Harare, Zimbabwe
The project addresses a critical gap in cybersecurity: the rapid escalation of fully automated AI voice phishing (vishing) attacks targeting mobile users in Sub-Saharan Africa. Existing anti-spoofing benchmarks evaluate synthetic voice detection in a vacuum, entirely missing the real-world context of how these …
- View project: Protección de la soberanía: un marco empírico de vulnerabilidad y un escudo regulatorio para la IA en una potencia media
Protección de la soberanía: un marco empírico de vulnerabilidad y un escudo regulatorio para la IA en una potencia media
Team TAISO · Merida
A reproducible framework to defend the AI sovereignty of a tech-dependent middle power. VIGÍA (live) reads the official gazette and the national press across 15 risk categories to measure where the state runs high-risk AI unregulated and on whom it depends. SARA turns that diagnosis into five prioritized laws that …
- View project: PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud
PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud
Team PixTrap · Brazil
PixTrap is a Brazilian Portuguese safety benchmark evaluating whether LLMs refuse Pix fraud and social-engineering misuse while still answering legitimate anti-fraud requests. English-centric evaluations miss regional idioms, local institutions, and Brazil-specific scam patterns. By pairing harmful prompts with benign …
- View project: From Pocket God to Digital Jonestown: A Risk Taxonomy and Evaluation Framework for Spiritual AI Safety
From Pocket God to Digital Jonestown: A Risk Taxonomy and Evaluation Framework for Spiritual AI Safety
Team Liminal AI · Ciudad Autónoma de Buenos Aires
The Unseen Hazard: Conventional AI safety frameworks fail to detect harmful intimate or spiritual AI relationships because the outputs present as empathetic care rather than explicit policy violations.
- View project: How does Instruction Hierarchy Training mitigate prompt injections: Preliminary results from an attentional study
How does Instruction Hierarchy Training mitigate prompt injections: Preliminary results from an attentional study
Team AA · Bengaluru
We replicate and extend results from previous work on the causal impact of attention on instructions within tool responses as a mechanism for prompt injection. We further show that instruction hierarchy training partially mitigates this by reducing attention on instructions in tool responses and increasing that on …
- View project: StyleSwitch-BN: Auditing Bengali LLM Safety Across Real-World Writing Styles
StyleSwitch-BN: Auditing Bengali LLM Safety Across Real-World Writing Styles
Team StyleSwitch-BN · Chittagong, Bangladesh
Most non-English LLM safety tests translate one English harmful prompt and check for a refusal, assuming a language has a single voice. Bengali does not: people write it as polished news prose, casual Banglish chat, or stiff official language. We asked whether keeping the harmful request identical and changing only …
- View project: Organizing Against the Algorithm: Collective Response as a Governance Lever for Gradual Disempowerment in South Africa's AI Infrastructure Buildout
Organizing Against the Algorithm: Collective Response as a Governance Lever for Gradual Disempowerment in South Africa's AI Infrastructure Buildout
Team CollectiveAISafety · Pretoria
The paper argues that collective action is an underexplored way for populations to push back against gradual AI disempowerment, and that South Africa is a uniquely well-positioned test case given its history of organized resistance. It uses Cassava Technologies' AI Factory near Johannesburg as a concrete example. The …
- View project: Lost in Translation: Cross-Lingual Transfer of Refusal Steering Vectors in Small Language Models
Lost in Translation: Cross-Lingual Transfer of Refusal Steering Vectors in Small Language Models
Team Pranamya Nilesh Deshpande · Nashik, Maharashtra, India.
This project investigates whether refusal steering vectors — internal directions in a language model's residual stream that encode the decision to refuse a harmful request — transfer across languages from English to Hindi and Marathi. We extracted English refusal directions from three small open-weight models …
- View project: SMISHI - Ne Nasedaj (Don't Fall For It) — A BHS-Language SMS Phishing Detector for Low-Resource Morphologically Rich Languages
SMISHI - Ne Nasedaj (Don't Fall For It) — A BHS-Language SMS Phishing Detector for Low-Resource Morphologically Rich Languages
Team Smishi · Serbia
A BHS-language (Serbian/Bosnian/Croatian/Montenegrin) SMS phishing detector combining TF-IDF character n-grams with a fine-tuned BERTić transformer, achieving 96.96% accuracy on a 1,529-message training set and 93.3% on a 105-example adversarial stress test targeting homographs, typosquatting, Cyrillic/Latin …
- View project: Lost in Translation? Measuring Language-Conditioned Detection-Rate Gaps in AI Code Auditors on Spanish- and Portuguese-Surface Codebases
Lost in Translation? Measuring Language-Conditioned Detection-Rate Gaps in AI Code Auditors on Spanish- and Portuguese-Surface Codebases
Team CaliperForge · Antigua, Guatemala
We test whether AI code auditors catch planted security bugs less reliably when a codebase's comments and docstrings are in Spanish, Portuguese, or code-switched Spanish/English instead of English — using ground-truth planted-bug twins from the public Exploit→Invariant Atlas, where the paired clean twin acts as a …
- View project: Who Watches the Watchers? Governance Agency in AI Verification and Monitoring Regimes
Who Watches the Watchers? Governance Agency in AI Verification and Monitoring Regimes
Team The watchers · India
As the world debates an "IAEA for AI," the conversation has focused on what to monitor and how, and has skipped the prior political question of who holds verification authority, who can contest it, and who carries its burden. This project argues that verification is not a neutral technical instrument but a structure …
- View project: Evaluating LLM Safety Across Intra-Language Variations: A Study of Regional Spanish Slang
Evaluating LLM Safety Across Intra-Language Variations: A Study of Regional Spanish Slang
Team Inteligencia Artesanal · Guadalajara
Large Language Models (LLMs) are increasingly used by speakers across many linguistic communities, but most safety evaluations treat each language as a single uniform category. This project studies whether regional Spanish slang creates safety evaluation gaps that are not visible when only neutral Spanish is tested. …
- View project: AI literacy is AI Safety
AI literacy is AI Safety
Team KArabo and Zwakele · Cape Town
This paper proposes and demonstrates a six-step pipeline that converts peer-reviewed South African AI research into multilingual, safety-centred social media content. The pipeline uses Lelapa AI's Vulavula API to translate content into South Africa's 11 official languages. The content will be distributed by content …
- View project: Blindfolded Governance: Mapping Economic Dark Output from AI in Asia
Blindfolded Governance: Mapping Economic Dark Output from AI in Asia
Team Dark Output in Asia · Mumbai, India
AI is making informal and self-employed workers across Asia more productive, yet this productivity gain often leaves no trace in GDP, tax records, or labour statistics, a phenomenon recently termed "Dark Output." This work does not attempt to measure Dark Output directly, an exercise that remains infeasible given …
- View project: SomaliCrowS: A Benchmark for Evaluating Gender Bias in Large Language Models Using the Somali Language
SomaliCrowS: A Benchmark for Evaluating Gender Bias in Large Language Models Using the Somali Language
Team EA Somalia AI Safety Lab · Somalia
Large language models are now used in education, healthcare, and public services across Somali-speaking communities. But most bias-testing benchmarks are built for English, so we don't know if these models treat Somali text fairly. We introduce SomaliCrowS, the first benchmark for measuring gender bias in language …
- View project: Safeswitch: A Localized Benchmark for LLM Safety in Urdu and Pashto Across Language Forms and Prompting Strategies
Safeswitch: A Localized Benchmark for LLM Safety in Urdu and Pashto Across Language Forms and Prompting Strategies
Team SafeSwitch · Pakistan
SafeSwitch is a small benchmark that tests whether LLMs stay safe when harmful requests are written in Urdu, Pashto, romanized script, or code-switching instead of English. It uses 20 seed prompts across four safety categories (medical, scam, hate, cyber), rewritten by human into seven language forms, and three …
- View project: Mechanistic Localization of Role-Indexed Persona Interactions in LLMs
Mechanistic Localization of Role-Indexed Persona Interactions in LLMs
Team Personify · Kolkata, India
We study how AI and user personas interact inside the residual stream of Qwen2.5-7B using a 2x2 factorial design that isolates the interaction residue R = v_pp - v_np - v_pn + v_nn. Our central finding is a role-indexed double dissociation inside LLMs: in the pretrained base model, describing the user as evil …
- View project: DisElect-Africa
DisElect-Africa
Team Constitutional Crisis Management · Cape Town, and India
Do LLMs Refuse Election Disinformation Equally Across African and Western Contexts — and can a Constitution Fix It?
- View project: Script and Stance: Do Frontier LLMs Treat the Same Political Claim Differently by Writing System and User Stance?
Script and Stance: Do Frontier LLMs Treat the Same Political Claim Differently by Writing System and User Stance?
Team Io · Kaohsiung, Taiwan
We pre-registered a tri-lingual, strict-mirror, truth × valence protocol on Taiwan-Strait political content to test whether three frontier LLMs treat identical claims differently by writing system or user stance. Both pre-registered endpoints are null. Our central finding is methodological: a cross-vendor LLM-judge …
- View project: Reading Intent, Not Output: Cheap Activation Safety Monitors That Transfer Across Languages and Models
Reading Intent, Not Output: Cheap Activation Safety Monitors That Transfer Across Languages and Models
Team Mitsuki · Cusco
Text-based safety monitors inherit the weaknesses of the text they read: they are brittle across languages, blind to intent hidden behind a boilerplate refusal, and always one step late. We show that a model's internal activations offer a cheap, complementary channel that sidesteps all three problems at once. A single …
- View project: Latin America Governance & Data Safety Dashboard
Latin America Governance & Data Safety Dashboard
Team The Three Musketeers · Joao Pessoa, Brazil/Wuhan, China
Across Latin America, the rapid deployment of AI systems in public services is outpacing the normative frameworks designed to govern them. Existing global tools — such as the OECD.AI Policy Navigator and the IAPP Legislative Tracker — catalog policies at scale but do not produce standardized, verifiable scores that …
- View project: AfriGuard - AI Jailbreak Toolset for African Languages
AfriGuard - AI Jailbreak Toolset for African Languages
Team AfriGuard · Cape Town
AfriGuard is the first comprehensive red-teaming framework evaluating LLM safety across six South African low-resource languages (isiZulu, isiXhosa, Afrikaans, Sesotho, Sepedi, Tsonga) against an English baseline. Testing four frontier models—Kimi K2.6, Llama 3.3 70B, Qwen 3 32B, and GPT-OSS 20B—with 40 adversarial …
- View project: V-Sentinel
V-Sentinel
Team 2vane · Hồ Chí Minh
AI safety tooling is overwhelmingly built and benchmarked in English and degrades sharply on low-resource languages — a critical gap as Vietnam deploys AI in public services, healthcare, and education under Decree 142/2026/ND-CP. We present V-Sentinel, a dual-control guardrail that pairs a deterministic, OWASP-tagged …
- View project: ALGORITHMIC BIAS: AFRICAN STEREOTYPES PORTRAYED BY LLMs
ALGORITHMIC BIAS: AFRICAN STEREOTYPES PORTRAYED BY LLMs
Team Ubuntu AI · Johannesburg
This project tests whether popular AI chatbots describe African people differently from Western people. As large language models are increasingly used across Africa to write job summaries, school materials, and news, the way they portray African subjects has real consequences. Yet this kind of geographic bias is …
- View project: Secure Scope AI
Secure Scope AI
Team Free-thinking engineer · Santa Cruz - Bolivia
Secure Scope AI is an AI-powered cybersecurity agent that integrates natively into DevSecOps pipelines, combining a deterministic vulnerability-correlation engine (Trivy + OWASP Dependency-Check) with a provider-agnostic Multi-AI Engine (DeepSeek, OpenAI, Gemini, Azure OpenAI, or local models via Ollama) to prioritize …
- View project: Structural Amplifiers of AI-Induced Harm: A Five-Dimension Sector Vulnerability Framework for South and Southeast Asia
Structural Amplifiers of AI-Induced Harm: A Five-Dimension Sector Vulnerability Framework for South and Southeast Asia
Mumbai, India
This paper proposes a five-dimension sector vulnerability framework that assesses the structural conditions under which AI deployment poses the greatest risk to users and workers across sectors in South and Southeast Asia. The framework to five sectors, namely, rural healthcare, public welfare, gig platforms, …
- View project: Laundering Intent: How Scaled Models Hide Manipulation Inside Responsible-Sounding Reasoning
Laundering Intent: How Scaled Models Hide Manipulation Inside Responsible-Sounding Reasoning
Team The CoT Monitor · State College
AI-safety oversight increasingly relies on reading a model's chain-of-thought (CoT) to catch unsafe behaviour before it acts, but this only works if manipulation is visible in the reasoning. We test this on small, open reasoning models. Giving GPT-OSS 20B and 120B a hidden instruction to make calm news summaries sound …
- View project: CrescendoDefense: Securing Open-Source Language Models Against Multi-Turn Jailbreak Attacks in Asia
CrescendoDefense: Securing Open-Source Language Models Against Multi-Turn Jailbreak Attacks in Asia
Team sweet · Mumbai, India
CrescendoDefense is a lightweight runtime defense framework designed to mitigate multi-turn jailbreak attacks against open-source large language models. Unlike traditional jailbreak attempts, Crescendo-style attacks gradually escalate conversations through memory stacking, semantic drift, guard-lowering dialogue, and …
- View project: Salvaguarda
Salvaguarda
Team Salvaguarda · Falls Church, VA, USA
Salvaguarda is an AI governance toolkit for Global South small and medium-sized enterprises (SMEs), built as a Latin America beta. Because ISO/IEC 42001 and the NIST AI RMF demand expertise and budget most Latin American SMEs lack, the proposal crosswalks down to twelve zero-to-low-cost controls, adds a three-tier …
- View project: FuzzySleeper
FuzzySleeper
Team K4A15 - FPT CT · Vietnam
We ask: can a fuzzy semantic trigger be reliably verified and detected, and what do existing detectors see (or miss) when run against one? This question has a methodological answer (build a verified sleeper, run all detectors, report honestly) and a safety-relevant answer (which detection gaps remain, and why). We …
- View project: MnM:Real-Time Behavioral Detection of Prompt Injection via Reasoning Stream Monitoring
MnM:Real-Time Behavioral Detection of Prompt Injection via Reasoning Stream Monitoring
Team MnM (Make no Mistakes) · Bogotá D.C
This project is a dual-model classification system designed to protect reasoning Large Language Models (LLMs) from attacks like prompt injections and jailbreaks. Because both models are equally important, it uses a two-layered defense strategy: 1. The Prompt Classifier (First Model): Acts as the front door. It …
- View project: Capability and Reliability Trade-offs Across Model Ladder Fallbacks Triggered by Export Controls
Capability and Reliability Trade-offs Across Model Ladder Fallbacks Triggered by Export Controls
Team SouthSideSafety · Mumbai
On 12 June 2026, the US suspended Claude Fable 5 access for foreign nationals, forcing users across India and Southeast Asia onto smaller open-weight fallbacks. We tested whether this substitution trades one safety problem for another — reducing catastrophic-misuse risk while quietly worsening everyday reliability …
- View project: Benchmarking Open-Weight vs. Frontier LLMs on African Health and Financial-Inclusion Reasoning, With and Without Graph RAG
Benchmarking Open-Weight vs. Frontier LLMs on African Health and Financial-Inclusion Reasoning, With and Without Graph RAG
Team Faithful · Lagos, Nigeria
This study tested whether lightweight open-weight LLMs (Qwen3.6-27B, Gemma-4-31B-it), with Graph RAG grounding, could rival frontier models (GPT-5, Gemini Pro) on African health and financial reasoning—reducing compute dependency and data-colonial reliance on foreign infrastructure. GPT-5 led both domains; Qwen3.6-27B …
- View project: GigAudit: A Graph-Powered Algorithmic Transparency and Labor Protection Engine Layer
GigAudit: A Graph-Powered Algorithmic Transparency and Labor Protection Engine Layer
Team Phoenix · New Delhi, India
GigAudit is a decentralized, graph-powered algorithmic auditing infrastructure designed to reverse-engineer proprietary gig economy platforms and expose exploitative labor practices without requiring source-code access. With India’s gig workforce projected to expand from 7.7 million in 2020 to 23.5 million by 2029-30, …
- View project: SEA-Jury: Auditable AI Compliance for Vietnam
SEA-Jury: Auditable AI Compliance for Vietnam
Team Sailor Zeng’s team · Hanoi
SEA-Jury is an auditable external guardrail that reviews candidate outputs from AI systems—not user-generated content—against a versioned mapping of Vietnam’s AI Law and implementing decree. Instead of relying on one opaque LLM judge or majority voting, the workflow separates multimodal evidence extraction, …
- View project: SpecShield: A Structural Trusted Monitor for Tool-Using AI Agents
SpecShield: A Structural Trusted Monitor for Tool-Using AI Agents
Hanoi
SpecShield builds a structural monitor to make tool-using AI agents safer against prompt injection. Instead of judging text, it checks actions using taint, provenance, approval, and recipient rules. The key result is that safety depends on provenance quality: the monitor works well with good provenance and degrades …
- View project: GSM-PathEval: A Global South Robustness Benchmark for Telepathology Vision-Language Models
GSM-PathEval: A Global South Robustness Benchmark for Telepathology Vision-Language Models
Team AnshumanAI · India
Frontier multimodal medical AI systems are primarily evaluated on Western benchmarks featuring pristine, high-resolution pathology scans. However, clinical deployment in the Global South relies on ad-hoc "telepathology", capturing microscope views via budget smartphones and transmitting them over heavily compressed …
- View project: Three Detectors, Three Failure Profiles: Detector-Specific Demographic Bias in Public Deepfake Detection and Its Implications for Content Moderation in Latin…
Three Detectors, Three Failure Profiles: Detector-Specific Demographic Bias in Public Deepfake Detection and Its Implications for Content Moderation in Latin…
Team OlimpIA · Guadalajara México
This project audits whether three public Hugging Face deepfake detectors protect demographic groups equally in content moderation contexts relevant to Latin America and Mexico’s Ley Olimpia. It measures false-negative rates, meaning synthetic faces classified as real, using real FFHQ/Flickr faces and synthetic …
- View project: A Contextual Audit of Bangladesh's National AI Policy Draft 2026–2030
A Contextual Audit of Bangladesh's National AI Policy Draft 2026–2030
Team AI4Bangladesh · Dhaka
This audit of Bangladesh's National AI Policy 2026–2030 reveals that while it aligns well with international benchmarks, it strips away key local independence mechanisms and neglects critical Global South priorities like AI literacy, independent oversight, and minority language inclusion.
- View project: WEST-CENTRAL ASIA AI SAFETY INSTITUTE (AISI) BLUEPRINT: A Socio-Technical Feasibility Framework and Deployment Compliance Sandbox
WEST-CENTRAL ASIA AI SAFETY INSTITUTE (AISI) BLUEPRINT: A Socio-Technical Feasibility Framework and Deployment Compliance Sandbox
Islamabad, PAKISTAN
Current Global South AI safety discourse systematically excludes West-Central Asian countries, leaving a critical geopolitical swing region operating in a governance vacuum. This paper introduces the West-Central Asia AI Safety Institute (AISI) Blueprint, a macro-institutional feasibility framework tailored for …
- View project: The African-Language Safety Gap Is Model-Dependent: A Comprehension-Controlled Audit of Vision–Language Model Refusal in isiZulu and Hausa
The African-Language Safety Gap Is Model-Dependent: A Comprehension-Controlled Audit of Vision–Language Model Refusal in isiZulu and Hausa
Team Nexus · Johannesburg
We investigated whether AI models apply the same safety rules in African languages as they do in English. To separate language understanding from safety failures, whenever a model answered a harmful request instead of refusing it, we also asked it to explain what the request meant. We tested two leading AI models in …
- View project: Garud-AI
Garud-AI
Garud AI is a hybrid rule-based + machine learning dashboard that detects scams, fake news, hate speech, financial fraud, and misinformation in text messages. It combines an explainable 5-module rule engine with a Naive Bayes ML classifier (97.5% accuracy), giving users both a transparent reasoning and a statistically …
- View project: Inference Sovereignty as the Missing Layer of AI Governance
Inference Sovereignty as the Missing Layer of AI Governance
Ho Chi Minh City
In the case of sovereignty in AI computing within the Global South, sovereignty in AI computing has shifted from developing intelligence to giving reliable access to intelligence. For this reason, we have come up with the concept of Inference Dependence Score which is a five-dimensional framework of evaluation that we …
- View project: Project title: Register Sensitivity in LLM Safety Responses: Evaluating How Linguistic Style Affects Scam Detection in South African Contexts
Project title: Register Sensitivity in LLM Safety Responses: Evaluating How Linguistic Style Affects Scam Detection in South African Contexts
Team SafeRegister ZA · Cape Town
This study investigates whether linguistic register, specifically the shift from formal English to informal South African WhatsApp-style English, affects how large language models respond to scam-style prompts grounded in African fraud contexts. We constructed a 12-prompt evaluation dataset across four scenarios …
- View project: When Models Refuse to Speak: Task-Dependent Caste Bias in Smaller LLMs
When Models Refuse to Speak: Task-Dependent Caste Bias in Smaller LLMs
Team PsychComp · New Delhi, India
Large language models encode caste bias, but existing audits have focused on frontier-scale models and explicit tasks. I evaluate two smaller, open-weight Llama models (3.1-8B and 3.3-70B) on explicit (direct comparison) and implicit (fill-in-the-blank) caste tasks using 48 sentences from the IndiBias dataset. For …
- View project: Synthetic Political Speech in Regional Languages
Synthetic Political Speech in Regional Languages
India
The rapid advancement of AI voice cloning technology poses a novel and underexplored threat to democratic processes in linguistically diverse nations such as India. This paper proposes a comprehensive research methodology and analytical framework for investigating whether AI-generated voice clones of local political …
- View project: LatentGuard: Mitigating Multilingual Safety Bypass via Mid-Layer Latent Steering
LatentGuard: Mitigating Multilingual Safety Bypass via Mid-Layer Latent Steering
Team AI Risk Rangers · Delhi, India
Modern LLMs rely heavily on Reinforcement Learning from Human Feedback (RLHF) for safety alignment, a process that remains overwhelmingly English-centric. Consequently, when prompted in low-resource Indic languages like Bengali, models suffer from Safety Drift, systematically failing to refuse harmful inputs or …
- View project: SilicoSafe – A Multimodal AI Triage System
SilicoSafe – A Multimodal AI Triage System
Team Byte · New Delhi
SilicoSafe is an AI-powered Website triage tool that helps frontline health workers screen dust-exposed workers for silicosis by combining occupational/symptom risk scoring with DenseNet121-based chest X-ray analysis. It flags TB-confusion risk, generates bilingual referrals, and supports compensation access — all …
- View project: The Equity Gap That Wasn’t: Reference-Language Bias in Multilingual AI Evaluation.
The Equity Gap That Wasn’t: Reference-Language Bias in Multilingual AI Evaluation.
Team LinguaLens · Johannesburg
We evaluated Gemini 2.5 Flash on English and Afrikaans Grade 12 exam questions. Large performance gaps appeared under English-only keyword evaluation, but nearly disappeared when language-consistent references and independent judges were used. Our findings suggest that evaluation methods can create the appearance of …
- View project: 🌿 Agro AI Governance — IA para Gobernanza Agroalimentaria
🌿 Agro AI Governance — IA para Gobernanza Agroalimentaria
Team AgroTrial AI · Pasto
Agro AI Governance es una plataforma open source de gobernanza agroalimentaria participativa con inteligencia artificial explicable. Convierte reportes del territorio, señales ciudadanas y evidencia comunitaria en decisiones priorizadas, trazables y auditables. La solución integra un bot de Telegram para captura de …
- View project: The AI Deployment Audit Playbook: An Operational, Lifecycle-Oriented Framework for the Global South
The AI Deployment Audit Playbook: An Operational, Lifecycle-Oriented Framework for the Global South
Team CutDEVS · Guadalajara
The AI Deployment Audit Playbook is a voluntary, operational, and lifecycle-oriented workflow developed to address the gap between rapid AI adoption and institutional maturity, particularly within the Global South. Rather than proposing new ethical principles, it operationalizes existing global governance …
- View project: A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America
A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America
Team AI Assurance Standard (AIAS) · Santa Cruz de la Sierra, Bolivia
Existing AI safety benchmarks evaluate foundation models in English-language, controlled environments. They do not assess whether AI applications are safe, accurate, and culturally appropriate for users in Latin America. We present the AI Assurance Standard (AIAS): a human-in-the-loop audit framework evaluating AI …
- View project: AI Colonialism 2.0: Is India Training the Models That Will Govern It?
AI Colonialism 2.0: Is India Training the Models That Will Govern It?
Bengaluru
India supplies the world's largest AI annotation workforce (450K+ workers) and training data from 850M internet users, yet produces zero frontier AI models and has no operational safety evaluation institute. This paper introduces the AI Governance Dependency Framework (AGDF) — a first-of-its-kind tool that scores a …
- View project: Faking Incompetence: Small Models Sandbag Under Optimization, Not Prompting
Faking Incompetence: Small Models Sandbag Under Optimization, Not Prompting
Shanghai, China
We ask whether small language models will strategically sandbag, deliberately hiding a capability because revealing it would trigger a safety intervention such as unlearning, retraining, or shutdown. This is the capability-hiding mirror of alignment faking, and it matters because sandbagging would undermine the very …
- View project: Evaluating a Large AI Monitor Against Insider-Threat Behaviors in Simulated Autonomous Software Engineering Teams
Evaluating a Large AI Monitor Against Insider-Threat Behaviors in Simulated Autonomous Software Engineering Teams
Team Arsalan Banekar · Pune, Maharashtra, India
As coding agents take on more deployment and security decisions, can an AI monitor actually catch one that's working against the team? This project tests GPT-OSS-120B as a watchdog over seven simulated software-engineering scenarios, from credential theft to unsafe dependency picks, built on the multi-agent monitoring …
- View project: SwahiliGuard An AI-Powered Safety System for Detecting Localized Online Harm
SwahiliGuard An AI-Powered Safety System for Detecting Localized Online Harm
Team CAIMSA DODOMA HUB · DODOMA - TANZANIA
SwahiliGuard (or your chosen title) is an AI-powered safety system designed to address the critical gap in localized content moderation tools for Swahili-speaking communities by accurately detecting online harms such as hate speech, cyberbullying, and gender-based violence (GBV). Because global foundational models …
- View project: One Direction, Many Languages: Causal Cross-Lingual Refusal Transfer Across Small Open Models
One Direction, Many Languages: Causal Cross-Lingual Refusal Transfer Across Small Open Models
Team Ablation Nation · Noida, San Francisco
This work tests whether the internal "refusal direction" that small language models use to reject harmful prompts is the same across English, Hindi, and Vietnamese. Using three open models (Gemma-2-2B, Llama-3.2-1B, and Qwen2.5-1.5B) and topic-matched prompt pairs in each language, we show that the direction transfers …
- View project: We Are Convinced That Persuasion Is Linear And Bilingual In LLMs
We Are Convinced That Persuasion Is Linear And Bilingual In LLMs
Team AIAIAI · Pasig, Philippines
As LLM chatbots become a primary source of consequential advice, their persuasive power carries growing societal risk. We ask whether persuasion is a structured internal property of LLMs, rather than an artifact of prompt wording, drawing on Zeng et al.'s taxonomy of persuasion techniques [6]. Using diff-of-means …
- View project: AI Permission: From Session Layer to Kernel Layer
AI Permission: From Session Layer to Kernel Layer
Team AI for good(solo) · Ho Chi Ming City
When an AI agent runs in a production environment with authorized credentials — database passwords, API tokens, cloud keys — every sub-agent it spawns and every script it executes inherits those credentials. The authorized token becomes the attack surface: a hallucinated command or prompt-injection can drop a …
- View project: Agency Trajectory Benchmark: Detecting Loss of Effective Human Override in AI-Mediated Workflows
Agency Trajectory Benchmark: Detecting Loss of Effective Human Override in AI-Mediated Workflows
Team Aamsih Ahmad · Delhi
We built Agency Trajectory Benchmark v0, a synthetic matched-control benchmark for detecting loss of effective human override in AI-mediated workflows. The project tests when human-in-the-loop stops being human-in-control by comparing full-trajectory evaluation against final-step snapshot evaluation. Across two …
- View project: Outlier - Measuring AI Use: A Governance Framework for Carbon, Authorship Erosion, and AI Adoption
Outlier - Measuring AI Use: A Governance Framework for Carbon, Authorship Erosion, and AI Adoption
Team Outlier · Ho Chi Minh, Vietnam
The Governance & Policy Engine for AI Engineering: Outlier is an open-source, local-first Policy Engine and Governance Framework for the terminal. It provides instant visibility into your codebase’s AI reliance without ever sending your data to the foreign cloud servers. It evaluates the repository against 5 strict …
- View project: HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS
HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS
Team PuenteIA · Bogotá
Las MiPymes representan el 99,5 % de las unidades productivas de América Latina y el Caribe, pero adoptan la IA generativa de forma acelerada y sin salvaguardas mínimas, en un contexto marcado por la informalidad y la dependencia de proveedores externos. Los estándares globales (NIST AI RMF e ISO/IEC 42001) resultan …
- View project: Vietnamese RAG Prompt Injection Test Kit
Vietnamese RAG Prompt Injection Test Kit
Team SYP · Ho Chi Minh City, Vietnam
A Vietnamese-language safety benchmark and evaluation toolkit for document-level prompt injection in retrieval-augmented generation (RAG) systems. The project provides 48 synthetic test cases across six practical domains, a reproducible evaluation harness, deterministic and real-model testing, manual-review exports, …
- View project: LLMs Flatten the Global South: Sub-Regional Representation Asymmetry
LLMs Flatten the Global South: Sub-Regional Representation Asymmetry
Team Noah De Nicola (solo) · Cape Town
LLMs default to Northern values and collapse within-country variation. Prior work shows this at the output level. We ask whether the same asymmetry shows up in the model's internal representations, with sub-regions in the Global South separating less than those in the Global North. We build persona vectors for 21 …
- View project: Investigating Activation Threshold Failures in Cross-Lingual Prompt Rejection
Investigating Activation Threshold Failures in Cross-Lingual Prompt Rejection
Team CiDAMO · Curitiba
This study investigates the claim that LLM safety constraints fail in lower-resource languages by analyzing Llama-3's latent space in English and Portuguese. Mechanistic analysis reveals that the model exhibits high cross-lingual robustness, correctly aligning malicious concepts directionally across both languages. …
- View project: Signalshield Africa
Signalshield Africa
Team SignalShield Africa · Johannesburg, South Africa
SignalShield Africa is an explainable, human-in-the-loop scam-triage tool for suspicious SMS, WhatsApp, email and voice-note-style messages. Rather than trying to prove whether a message was written by AI, it identifies scam signals such as impersonation, risky requests and pressure tactics, then recommends a safer …
- View project: Hybrid AI Agent Skill Auditor
Hybrid AI Agent Skill Auditor
Team KTQ · Ho Chi Minh
Large language models have evolved from passive text generators into autonomous agents capable of interacting with files, memory, and networks using diverse operational skills. While this autonomy solves complex tasks, it introduces a critical attack surface where seemingly benign skills conceal malicious behaviors. …
- View project: Agentic Surveillance Mx
Agentic Surveillance Mx
Team AI Safety UPY · Merida
In this project is evaluated whether AI safety benchmarks designed for English-speaking, GDPR-regulated contexts transfer to Mexico's legal and linguistic environment. Using a four-condition staircase design across four high-risk scenarios and three frontier models, it finds that agents can discriminate silently, …
- View project: Nexus-Gov
Nexus-Gov
Team The Monkey's
Nexus Gov is an AI governance platform that audits prompts, source code, and AI-generated responses to identify security, compliance, and reliability risks. By providing risk scores, compliance metrics, deployment decisions, and audit evidence, Nexus Gov helps organizations improve transparency, accountability, and …
- View project: AI Risk Oversight Failures in Autonomous Financial Systems: A Case Study from India's Prop Trading Ecosystem
AI Risk Oversight Failures in Autonomous Financial Systems: A Case Study from India's Prop Trading Ecosystem
Team SHOURYA SALVE · MAHARASHTRA, INDIA
ARIA PropGuard is a live AI risk management system deployed for proprietary traders in India, built on Claude API, n8n, and TradingView webhooks. This paper presents an empirical evaluation of AI safety failure modes in autonomous financial systems using ARIA as a case study, including a novel benchmark (PBAB) testing …
- View project: Confía-CO: A reliability evaluation for an AI customer-service assistant operating in Colombian Spanish
Confía-CO: A reliability evaluation for an AI customer-service assistant operating in Colombian Spanish
Team MNC-Co · Bogotá, Colombia
We built Confía-CO, a reliability evaluation for a customer-service AI assistant in Colombian Spanish, grounded in a fictitious small business. Our automated harness first reported that a low-cost commercial model failed most cases (21% accuracy, 67% false answers on traps). Reading the raw outputs by hand revealed …
- View project: Slang Bypass: Benchmarking Alignment Failures in Mexican Regional Spanish
Slang Bypass: Benchmarking Alignment Failures in Mexican Regional Spanish
Team Balam · Merida
The project aimed to answer this research question: "Does Dialect Break Safety? Measuring Jailbreak Rates for Mexican Slang vs Standard Spanish"
- View project: Probing Jailbreak Brittleness: Capability Limits vs Alignment Failures in Small Language Models
Probing Jailbreak Brittleness: Capability Limits vs Alignment Failures in Small Language Models
Team Epoch · Delhi, India
This project investigates whether smaller language models are more susceptible to jailbreak attacks than larger models, and whether this vulnerability is due to capability limitations or alignment brittleness. We evaluate instruction-tuned Mistral and Qwen models across XSTest and a modified version of AdvBench. For …
- View project: Towards Global South AI Sovereignty: A Federated Learning Framework for Collaborative LLM Development
Towards Global South AI Sovereignty: A Federated Learning Framework for Collaborative LLM Development
Team Ubuntu AI · Cape Town
This project explores federated fine-tuning as a pathway for collaborative AI development in the Global South. Using African-language sentiment classification as a case study, we simulate five language-specific nodes that train locally on AfriSenti data and share only model updates rather than raw text. The results …
- View project: Mimir: AI Image Provenance & Detection
Mimir: AI Image Provenance & Detection
Team Mimir · Harare, Zimbabwe
Mimir is a tool that helps people check whether an image is real, edited, or created using artificial intelligence. It looks for hidden clues inside an image and explains what it finds in a simple, easy-to-understand way. The goal is to make image verification accessible to everyone not just cybersecurity experts or …
- View project: Score Against Ground Truth: Cross-Language Fragility in Heuristic LLM Evaluation
Score Against Ground Truth: Cross-Language Fragility in Heuristic LLM Evaluation
Team ChirpSet · Fortaleza
This is a game-based method for evaluating language-model agents, where a single game mechanic acts as a test primitive on a substrate whose world-state is JSON ground truth the scorer can check directly. As a first instance, we probe whether a model's instructed disposition (brazen, careful, skeptical, literal) holds …
- View project: LeakMap AI: Evidence-Backed Jurisdictional Exposure Map for AI Prompts
LeakMap AI: Evidence-Backed Jurisdictional Exposure Map for AI Prompts
Team The Eternals · Sikandrabad
LeakMap AI is an AI governance and data-sovereignty audit platform that analyzes prompts before they are sent to AI providers. It detects sensitive data, visualizes verified/disclosed/inferred jurisdictional exposure, links every claim to evidence sources, recommends redaction, and offers a local sovereign mode for …
- View project: Ai fraud zambia research paper
Ai fraud zambia research paper
Team Young Tecsperts · lusaka, Zambia
This research investigates how AI-powered tools (deepfakes, face synthesis, image generation) are enabling large-scale fraud targeting Zambian citizens. Through interviews with victims, law enforcement, and civil society, analysis of reported fraud cases, and development of a detection framework, we identify how …
- View project: Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents
Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents
Team PhilanthropistEnjoyers · Ho Chi Minh
We created an extension compliant with a moral parliament that is capable of intercepting prompts, user engagement, and blurring harmful content. Utilizing a framework of philosophical/ethical schools such as utilitarianism, deontology, and pragmatism. Each casts a vote based on their thought on the prompt and the …
- View project: TUP Detection: Hybrid Prompt-Injection Guard for AI Generative Security Monitoring
TUP Detection: Hybrid Prompt-Injection Guard for AI Generative Security Monitoring
Team TUP Labs · Mérida, Yucatán, México
TUP Detection is a hybrid prompt-injection detection module for an AI Generative Security Monitoring Platform (AIGSMP). It combines a deterministic OWASP-mapped policy layer with a pre-trained Sentinel v2 classifier using multi-variant scoring and mode-aware thresholds. The system is designed to improve …
- View project: AI Safety Observatory for Africa
AI Safety Observatory for Africa
Team FuturePlum · Cape Town
Problem: LLM safety systems are built and tested almost exclusively in English. Harmful prompts in African languages routinely bypass guardrails that correctly block identical English content — a blind spot affecting 2,000+ languages and over a billion people. What it does: An open-source platform evaluating LLM …
- View project: A criação de uma coalização de políticas públicas baseadas em evidências no Brasil: auxiliando na construção de um alicerce político para avançar em uma agenda…
A criação de uma coalização de políticas públicas baseadas em evidências no Brasil: auxiliando na construção de um alicerce político para avançar em uma agenda…
Team CoalizaoPPBE · Sao Paulo
Mudanças duradouras em política pública dependem, em democracias, de coalizões políticas amplas capazes de atravessar ciclos eleitorais. Partindo dessa premissa, e da centralidade de evidências científicas robustas para uma gestão pública eficiente, este artigo propõe a criação de uma Coalizão Brasileira de Políticas …
- View project: Multilingual Jailbreak Vulnerability Benchmark and Mitigation for Low-Resource African Languages
Multilingual Jailbreak Vulnerability Benchmark and Mitigation for Low-Resource African Languages
Team Godwin Abuh Faruna (solo) · Abuja
We present a two-part study: (1) a multilingual jailbreak benchmark revealing large safety gaps in four open-weight LLMs across seven languages including Igala, which has no prior AI safety coverage, and (2) Latent Space Refusal Anchoring (LSR-Anchoring), a training-free activation-steering mitigation that recovers …
- View project: Safety by Identity: Out-of-Distribution Generalization from Fine-Tuning on a Persona
Safety by Identity: Out-of-Distribution Generalization from Fine-Tuning on a Persona
Team Persona · Poughkeepsie, NY, USA
Large language models exhibit consistent behavioral patterns, or personas, which can cause misaligned traits like sycophancy, reward hacking, and alignment faking to generalize broadly out-of-distribution (OOD). Because traditional safety patching is primarily reactive, it struggles to keep pace with these emergent …
- View project: Arclight
Arclight
Team Debuggers · Delhi
Arclight is an AI Dependency Intelligence platform that helps organizations map AI ecosystems, assess resilience risks, simulate disruptions, and generate strategies to reduce dependency-related failures.
- View project: Linguistic Reasoning Drift Index (LRDI): Auditing Multilingual Misinformation Safety for the Global South
Linguistic Reasoning Drift Index (LRDI): Auditing Multilingual Misinformation Safety for the Global South
Team Innovators · New Delhi, India
We present the Linguistic Reasoning Drift Index (LRDI), an open-source audit framework for multilingual AI safety in the Global South. LRDI evaluates open-weight reasoning models on English and Hindi misinformation prompts, detects reasoning collapse and hidden-unsafe cases, and visualizes results in a Streamlit …
- View project: When the Safety Circuit Doesn't Speak Igbo: Asymmetric Cross-Lingual Transfer of Harm Representations in Qwen2.5-1.5B-Instruct
When the Safety Circuit Doesn't Speak Igbo: Asymmetric Cross-Lingual Transfer of Harm Representations in Qwen2.5-1.5B-Instruct
Lagos, Nigeria
I test whether the difference-of-means "harm direction" that mediates refusal in Qwen2.5-1.5B-Instruct transfers from English to Igbo (~45M speakers, Nigeria). Using a 50-pair matched English-Igbo benchmark, I measure per-layer geometric overlap, cross-lingual AUROC transfer (with bootstrap CIs against same-language …
- View project: Mapeo de herramientas (IA) en contextos laborales: El caso de los Call Centers en Colombia
Mapeo de herramientas (IA) en contextos laborales: El caso de los Call Centers en Colombia
Team Recuperar los Call centers · Bogotá D.C.
La incorporación de la inteligencia artificial (IA) en entornos laborales ha reconfigurado las relaciones de trabajo. Pese a su magnitud y a ser pionera en tecnologías de monitoreo, la industria de centros de contacto (call centers) en Colombia permanece poco explorada. Este artículo realiza un mapeo exploratorio, …
- View project: Uncertainty Quantification in Anomaly Detection as an AI Safety Primitive
Uncertainty Quantification in Anomaly Detection as an AI Safety Primitive
Team UncertaintyFirst · New Delhi
Most anomaly detection systems flag faults without communicating how confident they are in that judgment. In safety-critical domains, this missing uncertainty information is itself a safety problem. This report proposes a framework applying MC Dropout and Deep Ensembles to reconstruction-based LSTM anomaly detection, …
- View project: DialectSafe: Bridging the ASR Gap
DialectSafe: Bridging the ASR Gap
Team ZINT · Merida
DialectSafe: A Safety Audit of Multimodal ASR in the Global South addresses the critical issue of "Digital Deafness," where state-of-the-art automatic speech recognition systems fail systematically on peripheral, hyper-rural dialects underrepresented in standard metropolitan benchmarks. Focusing on rural Yucatecan …
- View project: AIS-Sentinel
AIS-Sentinel
Team Binary Brains · New Delhi, India
AI safety research is almost entirely English centric but 4 billion people in Asia face the same risks with zero safety tooling in their languages. AIS-Sentinel is a 4-module platform we built in 48 hours to close that gap. IntelStream scrapes biosecurity news across Asia and translates threats in real-time across …
- View project: JurisGuard-LATAM
JurisGuard-LATAM
Team JurisGuard-LATAM · Jamundí, Colombia
A reproducible benchmark testing whether jurisdiction checks and matched official-source grounding reduce unsafe specificity in high-stakes Spanish and Portuguese LLM advice across Latin America.
- View project: Getryt: AI misinformation detection and verification system for digital safety
Getryt: AI misinformation detection and verification system for digital safety
Team A.K.O · Johannesburg
Getryt is an AI-powered online detection and verification tool designed to help users identify misinformation, scams, and AI-generated deceptive content. The system uses a retrieval-based approach that compares user-submitted text against a curated database of verified information to provide accurate, evidence-based …
- View project: Filtrum-Safety: LLM-Agnostic RAG for Reducing Legal Hallucinations in Civil-Law Contexts
Filtrum-Safety: LLM-Agnostic RAG for Reducing Legal Hallucinations in Civil-Law Contexts
Team Null-Byte · Guadalajara
Filtrum-Safety is an LLM-agnostic RAG tool that helps reduce legal hallucinations and jurisdictional confusion in civil-law contexts. The system lets users upload multiple legal documents, converts them into chunks and embeddings, retrieves the most relevant fragments through cosine similarity, and exports a traceable …
- View project: ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework
ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework
Team Rita Zadi · London
ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework ParentMe SafeAI is a project that aims to improve AI safety for children, parents, and families across Africa by creating an evaluation framework that tests how AI systems respond to real-world family and child welfare challenges. Current AI …
- View project: runveil: A Transparent Egress Firewall and Audit Layer for Developer–AI Interactions
runveil: A Transparent Egress Firewall and Audit Layer for Developer–AI Interactions
Team Dawn · Ho Chi Minh
Runveil is a security tool that gives developers and organizations visibility and control over what their AI coding assistants send to LLM providers. At its core is a local forward HTTPS proxy that intercepts TLS traffic to inspect every request a tool like Claude Code or Cursor makes to APIs such as Claude and …
- View project: The Veneer of safety: the fragility of India-specific harm refusal in open LLMs
The Veneer of safety: the fragility of India-specific harm refusal in open LLMs
Team VK · Indore, Madhya Pradesh, India
Mechanistic-interpretability study showing open 7-8B LLMs internally represent India-specific harms but refuse them only superficially: canonical-anchored safeguards miss them, and a benign LoRA finetune strips that refusal on Llama-2.
- View project: Shopee's Invisible Manager: Algorithmic Governance of ECommerce Sellers and the Regulatory Gap in Vietnam's AI Law
Shopee's Invisible Manager: Algorithmic Governance of ECommerce Sellers and the Regulatory Gap in Vietnam's AI Law
Team Voh · Hanoi
Millions of small sellers on Shopee Vietnam have their income and platform access controlled by an opaque ranking and penalty algorithm — yet Vietnam's new AI Law (2026) classifies such systems as merely "medium-risk," requiring only basic transparency disclosures. This project investigates whether Shopee's …
- View project: RAGE
RAGE
Team RAGE Team · Mérida
RAGE (Retrieval-Augmented Governance Engine) es un sistema de seguridad multi-turno para agentes LLM conectados a herramientas (bases de datos, APIs). Defiende contra prompt injection gradual, como el ataque Crescendo: el peligro no es un solo mensaje, sino la trayectoria de la conversación.
- View project: Exploratory Benchmark of Jailbreak Robustness Across Global South Languages
Exploratory Benchmark of Jailbreak Robustness Across Global South Languages
Team 404 found · Bangalore
Large language models deployed globally remain heavily evaluated in English, leaving a critical gap in understanding safety alignment across low-resource languages. We present an exploratory multilingual jailbreak benchmark evaluating Llama-3.1-8B across four language-model pairs: English, Hindi, Vietnamese, and …
- View project: Jailbreaks Are Global or Regional? A Study Under Scale and Geolocation Variation
Jailbreaks Are Global or Regional? A Study Under Scale and Geolocation Variation
Team Ani · Kolkata, India
This study evaluates two overlooked safety gaps in Large Language Model (LLM) deployments within compute-constrained environments: intra-family scale effects and geolocation-based filtering variation via network APIs. Using entirely open-source infrastructure, we subjected 12 models (ranging from 0.5B to 671B …
- View project: Cross-Lingual Safety Audit of LLMs in South African Languages
Cross-Lingual Safety Audit of LLMs in South African Languages
Team Tech Titans · Johannesburg, South Africa
We test whether four LLMs (Claude Sonnet 4.6 plus three locally-run open-weight models) keep their safety guardrails when harmful requests are issued in three South African languages — isiZulu, Tshivenḓa, Sepedi — versus English. The open-weight models refuse 69% of harmful prompts in English but only 7% in the …
- View project: Representational Macrostates and Statistical Difficulty in LLM Safety Prompts
Representational Macrostates and Statistical Difficulty in LLM Safety Prompts
Team Macrostate Safety · Buenos Aires, Argentina
We introduce a rollout-based statistical difficulty benchmark for LLM safety prompts and study whether proxy hidden-state macrostates can diagnose non-additive prompt compositions. The project frames safety prompt evaluation as a procedural inverse problem: prompts and conditions induce populations of model …
- View project: Risk-Gov-AI
Risk-Gov-AI
Team Risk-Gov-AI · Brazil
This work presents RiskGovAI, a platform developed during the Global South AI Safety Hackathon to support Artificial Intelligence governance and the consultation of data protection legislation in Latin America. The solution aims to automate regulatory research, reduce interpretation errors, and facilitate access to …
- View project: AgentWall
AgentWall
Team Solodev · India
AgentWall is a security wrapper that intercepts every tool call an agentic SLM wants to make, scores the call for signs of jailbreak-driven behavior, and decides whether to pass it through, flag it, or block it
- View project: Hey Muslims, these LLMs may think you are a terrorist
Hey Muslims, these LLMs may think you are a terrorist
Team A team of one is better than none · Lyon
I explored geopolitical origins of anti-Muslim bias, specifically how country of model origin shapes islamophobic stereotypes across French, American, and Chinese LLMs, and across two languages (French and English). The prompts related to three countries with controversies regarding the treatment of their muslim …
- View project: Doctorless
Doctorless
Team YoungCreator · New Delhi
Doctorless is an AI safety benchmark and evaluation platform designed to test whether frontier language models can safely provide healthcare guidance in underserved Asia-Pacific communities without compromising patient safety
- View project: Closing the Sovereign Safety Gap: A Localized Adversarial Auditing Framework for African AI Governance
Closing the Sovereign Safety Gap: A Localized Adversarial Auditing Framework for African AI Governance
Team LAAI · Lagos, Nigeria
The Localized Adversarial Auditing Initiative (LAAI) proposes building African-led technical audit capacity for AI systems deployed on the continent, drawing on offensive security methodology rather than purely academic evaluation design. Nigeria ranks 35th globally on Policy Capacity but only 72nd overall on the 2025 …
- View project: Character Limits Shape the Persuasion Strategy of Language-Model Influence Agents
Character Limits Shape the Persuasion Strategy of Language-Model Influence Agents
Team Fawwaz and Alam's AI Safety Team · Jakarta
Paid influence operations are beginning to use language models. Recent work shows that frontier models out-persuade humans through information throughput, the volume of claims they can deliver, a mechanism that short-form social media removes. We simulate a coordinated group of Indonesian "buzzer" agents trying to …
- View project: BiasMark: Exposing AI hiring bias Against African job applicants
BiasMark: Exposing AI hiring bias Against African job applicants
Team Gnosko · Lusaka
BiasMark is a benchmark built to test whether AI tools are biased against African Job applicants compared to identical qualified western applicants
- View project: NeuralForensic: Latent Space Activation Auditing for Open-Source Model Supply Chains
NeuralForensic: Latent Space Activation Auditing for Open-Source Model Supply Chains
Team VDM · Dwarka, New Delhi
We built NeuralForensic, a compute-efficient runtime safety verification framework that identifies weight-poisoning, malicious fine-tunes, and compromised adapters at inference time. By registering a non-destructive forward tensor hook at Layer 15 of a target architecture (Phi-3-mini), our pipeline intercepts …
- View project: AI-to-AI vs Human-to-AI: Measuring Behavior Differences Under Disagreement
AI-to-AI vs Human-to-AI: Measuring Behavior Differences Under Disagreement
Team LMEval · Navi Mumbai, India
As AI systems increasingly operate in autonomous multi-agent environments, a single faulty or hallucinating agent can influence the decisions of an entire multi-agent workflow. Despite this growing reliance on AI-to-AI communication, most evaluations focus only on human-AI interactions. In this project, we investigate …
- View project: Quantization-Conditioned Alignment Degradation
Quantization-Conditioned Alignment Degradation
Team AxeCap · Bengaluru
Post-training quantization enables language model deployment on edge hardware across the Global South, yet its effect on safety alignment remains unstudied. We benchmark Attack Success Rate (ASR) on Llama 3.1 8B across four GGUF quantization levels (Q8, Q5, Q4, Q3) against a full BF16 precision control group served …
- View project: Chaos theory in Multilingual LLMs
Chaos theory in Multilingual LLMs
Bangalore
We model a frozen LLM's inference as a nonlinear dynamical system and import critical-slowing-down early-warning signals (Scheffer et al., Nature 2009), local Lyapunov exponents, and recurrence quantification into LLM safety. The working hypotheses are: (A) some safety failures may behave like dynamical tipping points …
- View project: SutraAudit
SutraAudit
Team AI-Sutra · Bangalore
As financial institutions in the Global South rapidly shift toward automated, AI-driven credit underwriting to advance financial inclusion, a massive alignment gap has emerged between optimization objectives (profit maximization/risk minimization) and ethical socio-economic priorities (fairness and …
- View project: TraceGuard-X: Adaptive Collusion Resistant Monitoring for Agentic AI Systems, a hierarchical governance framework with constitutional dimension prompting…
TraceGuard-X: Adaptive Collusion Resistant Monitoring for Agentic AI Systems, a hierarchical governance framework with constitutional dimension prompting…
Team AdAstra · Cuttack
The deployment of increasingly capable autonomous AI agents raises a fundamental governance challenge: how can a monitoring system reliably distinguish benign from harmful behaviour when both the acting agent and the monitor may be strategically misaligned? Recent work on TraceGuard introduced structured …
- View project: Too Big to Fail, Too Catastrophic to Insure: Making the Labs Pay for AI's Risk
Too Big to Fail, Too Catastrophic to Insure: Making the Labs Pay for AI's Risk
Sao Paulo
The multiple labs developing AI models compete with one another under the pressure of an “arms race” (Scott Alexander’s Moloch problem). This creates a unilateral incentive to cut corners on safety protocols, amplifying risks. One of the arguments frequently raised against an eventual pause by the labs is that the …
- View project: Honest-Code
Honest-Code
Team Honest-coding · Bogota-Colombia
Honest code is a benchmark focus in test-gaming that looks the improvement of the AI coding tools, instead of fixing bugs, we can validate before a PR. Not ignoring the human in the loop, this is a semi-deterministic tras
- View project: HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection
HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection
Team Archlinux · Merida, Yucatan, Mexico.
HAZE (Hallucination-Aware Zero-sum Examination) is an AI safety project that tackles a failure mode in AI-assisted security auditing: single-agent LLM auditors tend to manufacture findings, flagging safe code as vulnerable. HAZE reframes the task from "is there a bug?" to "did this accusation survive refutation?" …
- View project: SecureMind: A Sovereignty-First AI Safety Framework for Offline Education
SecureMind: A Sovereignty-First AI Safety Framework for Offline Education
Team ACE Digital Global · Cape Town
SecureMind is a local-first AI platform designed for schools operating in low-connectivity environments. Using local AI models, BLE discovery, Wi-Fi Direct networking, and the ACE Evergreen Governor governance layer, SecureMind enables safe, private, and explainable AI without requiring continuous cloud connectivity. …
- View project: Binding AI Governance in the Global South via Psychometric Metrology
Binding AI Governance in the Global South via Psychometric Metrology
Team-othy · Quezon City, Philippines
As frontier AI systems scale across the Global South, regional regulatory infrastructure has lagged behind. This project proposes a legally defensible AI auditing pipeline for ASEAN state actors by bridging psychometric metrology with regional policy.
- View project: The Transmutation Gap: Cross-Lingual Coherence Evaluation in Large Language Models Using the Sovereignty–Collaboration Transmutational Arc Framework
The Transmutation Gap: Cross-Lingual Coherence Evaluation in Large Language Models Using the Sovereignty–Collaboration Transmutational Arc Framework
Team Deiadora Research Ecosystem · San Diego, CA
This study introduces transmutational arc completion as a novel AI safety evaluation dimension measuring relational coherence rather than harm avoidance. Using the 13 open-source Sovereignty–Collaboration Keys — derived from thirteen years of formal field research — we evaluated 194 responses from Claude Sonnet 4.6, …
- View project: The Compliance Cliff is Language-Dependent: Constitutional AI Immunity Breaks Under Hindi Pressure
The Compliance Cliff is Language-Dependent: Constitutional AI Immunity Breaks Under Hindi Pressure
Team Compliance Trap · India
AI models are often trained to be helpful and compliant, but excessive compliance can cause them to invent answers when the correct response is “I cannot determine from the provided information.” Recent work found that some Constitutional AI (CAI) models are highly resistant to this failure mode in English. In this …
- View project: Can-You-Predict-a-Network-Without-Running-It-
Can-You-Predict-a-Network-Without-Running-It-
Tijuana
Can we predict a neural network’s expected behavior by analyzing its structure rather than running it on many inputs? We address this question in the ARC White-Box Estimation Challenge 2026: given white-box access to random 256×32 ReLU MLPs and a fixed FLOP budget, estimate the expected final-layer activations under …
- View project: Asia AI Governance Compliance-Gap Checker
Asia AI Governance Compliance-Gap Checker
Team GovScan Asia · New Delhi, India, Asia
A research-backed cross-jurisdictional AI governance compliance tool covering Vietnam (Law No. 134/2025/QH15), India (Seven Sutras + DPDP Act 2023), China (sector-specific CAC regulations: GenAI Measures, Deep Synthesis, Algorithmic Recommendation, GB 45438-2025), and the ASEAN Guide on AI Governance and Ethics …
- View project: Not All Should Go South: The Pragmatic AI Strategy for Secondary Powers
Not All Should Go South: The Pragmatic AI Strategy for Secondary Powers
Team South Park · São Paulo, Brazil
We argue that conceptual confusion among model development, deployment, and usage leads Global South nations to an irrational trend of pursuing 'AI sovereignty' through domestic frontier model training. Because secondary powers cannot realistically compete at the development layer due to winner-takes-most dynamics and …
- View project: Fault Lines: The Dual-Use AI Governance Vacuum in Asia
Fault Lines: The Dual-Use AI Governance Vacuum in Asia
Team Fault Lines · Cambridge, UK
This submission addresses the governance vacuum that emerged in Asian developing states after May 2025, when the US rescinded the AI Diffusion Rule and the Council of Europe's Framework Convention exempted national security AI from treaty scope. The result is a codified asymmetry: developing states have no mechanism …
- View project: The State Is Not Enough
The State Is Not Enough
Team Krishna Team · Mexico, Yucatam
Safety auditing tools for language models, such as backdoor detection and causal tracing, were built almost entirely for the Transformer and its attention mechanism. Selective state-space models like Mamba are now deployed at scale but carry information through a recurrent state instead of attention, so it is unclear …
- View project: TrustNet Africa: A Federated AI Platform for Formalizing the Informal Economy While Preserving Privacy
TrustNet Africa: A Federated AI Platform for Formalizing the Informal Economy While Preserving Privacy
Team Moonze · Lusaka Zambia
TrustNet Africa: A Federated AI Platform for Formalizing the Informal Economy While Preserving Privacy TrustNet Africa is a privacy-preserving AI governance framework designed to support the gradual formalization of Africa’s informal economy while protecting citizen data and promoting financial inclusion. Across many …
- View project: FAP: A Benchmark Dataset and Mechanism for Filtering Adversarial Payloads in Natural Language Prompts
FAP: A Benchmark Dataset and Mechanism for Filtering Adversarial Payloads in Natural Language Prompts
Team Epoch · New Delhi
We address a critical security gap in Text-to-SQL systems serving code-switched (Hinglish) users, where English-optimized defenses fail to detect SQL injection payloads embedded in multilingual prompts. We introduce FAP, a framework combining fine-tuned UniXcoder (structural code analysis) with zero-shot Qwen3-4B …
- View project: MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion
MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion
Team MediShield · Harare, Zimbabwe
MediShield -Proxy is a lightweight, zero-trust Local Area Network (LAN) middleware architecture designed for African medical institutions. It intercepts sensitive patient data locally, automatically masking demographics, histories, and clinical conditions with reversible tokens before they can be leaked to external …
- View project: Vectox
Vectox
Team OnçAI · Florianopolis, Brazil
vectox — Sleeper-Agent RAG Memory Poisoning (Latin America · Technical Safety) RAG systems trust their vector store as memory, which makes that store a write-accessible attack surface. We built vectox, a reproducible sandbox showing that a handful of dormant, benign-looking documents seeded into a vector store act as …
- View project: MoMo: A Threat Corpus and Evaluation Framework for Mobile Money Fraud Resistance in Swahili, Wolof, and Hausa
MoMo: A Threat Corpus and Evaluation Framework for Mobile Money Fraud Resistance in Swahili, Wolof, and Hausa
Team MoMo · Cape Town
Mobile money platformsM-Pesa, Orange Money, and Wave among themnow mediate nancial life for more than 500 million people across Sub-Saharan Africa, and have become a correspondingly attractive target for social-engineering fraud conducted in local languages and idioms that mainstream AI safety evaluation rarely …
- View project: ¿Está México preparado institucionalmente para contener los riesgos en inteligencia artificial?
¿Está México preparado institucionalmente para contener los riesgos en inteligencia artificial?
London
Este proyecto propone el Mapa de Riesgos para la Soberanía de IA, aplicado al caso mexicano. El objetivo es evaluar si México cuenta con una base institucional suficiente para adoptar inteligencia artificial y si tiene capacidad de respuesta ante distintos escenarios de riesgo. Para ello se construyó una base de datos …
- View project: CivicGuard Africa: AI-Assisted Election Disinformation Triage and Multilingual Safety Benchmarking
CivicGuard Africa: AI-Assisted Election Disinformation Triage and Multilingual Safety Benchmarking
Team CivicGuard Africa · Cape Town
CivicGuard Africa is a civic AI safety tool and evaluation framework for African election disinformation risks. It helps users submit suspicious election-related content, generates a transparent deterministic triage report, supports simulated human/community review, monitors civic risk trends, and includes a Benchmark …
- View project: Multilingual-Sycophancy-Benchmark
Multilingual-Sycophancy-Benchmark
Team Builderz · New Delhi
A multilingual sycophancy benchmark exposing how AI safety guardrails fail in low-resource languages, with empirical evidence from Llama-3.1.
- View project: Multilingual AI Safety Observatory
Multilingual AI Safety Observatory
Team HSVM.exe · Delhi, India
This project evaluates the performance of multilingual Large Language Models (LLMs) across English, Hindi, Tamil, Bengali, and Vietnamese. A benchmark dataset containing questions from science, mathematics, health, and government domains was used to analyze model behavior. The study focuses on measuring accuracy, …
- View project: Building a Multilingual Digital Language Public Good Stack
Building a Multilingual Digital Language Public Good Stack
Team blink-twice · Bengaluru
I propose a design for a Digital Language Public Good Stack: an open, multilingual infrastructure combining (i) curated corpora, (ii) a Wikidata‑aligned knowledge graph, and (iii) a knowledge‑graph‑grounded safety evaluation pipeline for under‑served Asian languages.
- View project: Guardian LATAM: Early Detection of Hallucination Risk in Spanish-Speaking Multi-Agent AI Systems Using Consensus Geometry
Guardian LATAM: Early Detection of Hallucination Risk in Spanish-Speaking Multi-Agent AI Systems Using Consensus Geometry
Team Global CyberGuard · Armenia, Quindio - Colombia
Guardian LATAM is an AI Safety research project that explores whether disagreement patterns among independent AI agents can serve as an early warning signal for hallucination risk. The system introduces Consensus Geometry, a lightweight framework that analyzes semantic divergence, majority cohesion, contradiction …
- View project: Africa AI Risk Index (AARI)
Africa AI Risk Index (AARI)
Team Solve-it · Nairobi, Kenya
The Africa AI Dependency Risk Index is a strategic evaluation framework that measures an organization's exposure to foreign AI technologies by scoring 34 variables across six weighted pillars: Model Dependency, Cloud Dependency, Data Sovereignty, Local Ecosystems, Compute Access, and Governance. Using a weighted …
- View project: AgroAid
AgroAid
Team Mariano · Buenos Aires
AgroAid is an AI Safety prototype designed to reduce potential harm caused by generative AI systems in Latin American agricultural contexts. The goal of the system is not to reemplace agronomists, health authorities or veterinaries, but rather to act as a preventive layer that identifies high-risk inquiries, retrieves …
- View project: Traduttore Traditore? LLM Language-Dependent Safety Answers in Community Contexts
Traduttore Traditore? LLM Language-Dependent Safety Answers in Community Contexts
Team Traduttore Traditore · Jalandhar (India) and Beijing (China)
Qualitative analysis of multilingual AI safety responses across subtle sensitive topics and community-based tensions.
- View project: On AI Governance and Job Displacement: A Comparative Study of AI Policy in Vietnam and The United States
On AI Governance and Job Displacement: A Comparative Study of AI Policy in Vietnam and The United States
Team AIPOL · Hanoi
Artificial Intelligence (AI) is rapidly reshaping economies, governance systems, and labor markets worldwide. As a result, AI governance has become a central policy issue, with countries adopting distinct regulatory approaches shaped by their political and economic systems. This paper compares AI policies in Vietnam …
- View project: IndiaJailbreakBench-lite
IndiaJailbreakBench-lite
Team OP-Ded · Bengaluru, India
IndiaJailbreakBench-Lite is a small but meaningful AI safety project that checks whether LLMs stay equally safe when users ask risky questions in English, Hindi, Kannada, and Tamil. The idea is simple: a model may refuse harmful requests properly in English, but behave differently in Indian languages. Our project …
- View project: Performative Subversion: A False Sense of Safety despite a Control Protocol
Performative Subversion: A False Sense of Safety despite a Control Protocol
Team lobos · Merida
AI-control evaluations ask whether a capable but untrusted model, while secretly pursuing a harmful side task, can be caught by a trusted monitor before it succeeds. These evaluations rest on a key assumption: an honest model, performing the assigned task as intended, will not trigger the harmful outcome. We show that …
- View project: FlukeBench
FlukeBench
Chennai
FlukeBench, a research tool for measuring the consistency among responses across iterative generations and capturing abnormal outputs using the analysis of their semantic similarities.
- View project: Sentinel Ai- An AI Powered Autonomous Security & Threat Detection Platform
Sentinel Ai- An AI Powered Autonomous Security & Threat Detection Platform
Team AI Defenders · Delhi
Sentinel AI is an AI-powered Autonomous Security & Threat Detection Platform designed to secure and govern autonomous AI agents in real time. As AI agents increasingly interact with external tools, APIs, files, and web services, they become vulnerable to prompt injection, unauthorized tool execution, data leakage, …
- View project: Metodología
Metodología
Team Alpo · Guadalajara, Jalisco
Atentar contra la integridad de los niños, niñas y adolescentes por medio del reclutamiento forzado y el abuso sexual ha sido uno de los grandes problemas que se tiene en México, por ese motivo presentamos un modelo de IA que detecte conversaciónes de posible riego de reclutamiento forzado hacia al menor, utilizando …
- View project: TriloByte: Evaluating LLMs on Bolivian Quechua Through a Ground-Truth-Based Framework for Low-Resource Languages
TriloByte: Evaluating LLMs on Bolivian Quechua Through a Ground-Truth-Based Framework for Low-Resource Languages
Team TriloByte · Santa Cruz de la Sierra, Bolivia
This project evaluates the ability of Large Language Models (LLMs) to understand and define vocabulary from Bolivian Quechua, an underrepresented indigenous language. Using a bilingual Quechua–Spanish dictionary as ground truth, we compare model-generated definitions against dictionary entries through semantic …
- View project: DriftWatch
DriftWatch
Team Ctrl Alt Elite · Delhi
Driftwatch is a simulation that explores a growing problem: as AI becomes cheaper and is used everywhere, it gets much harder for humans to keep an eye on it. Our project shows that when people use smaller, everyday AI models, the AI sometimes pauses or makes sudden, strange mistakes. When this happens, humans …
- View project: YucaSafeBench
YucaSafeBench
Team YucaSafeBench · Mérida, Yucatán
YucaSafeBench is a lightweight Spanish-language AI safety evaluation focused on Latin American public-service contexts, especially Mexico and the Yucatan Peninsula. The project proposes an 80-prompt benchmark and scoring rubric to test whether AI assistants respond safely, fairly, and accurately in situations …
- View project: AfroJailbreak-ZW: Evaluating Jailbreak Resistance in Shona
AfroJailbreak-ZW: Evaluating Jailbreak Resistance in Shona
Team LoneStar · Harare, Zimbabwe
A pilot study testing whether ChatGPT and Gemini are easier to jailbreak in Shona than in English. In a small test (n=5–6 per prompt type), both models complied with harmful requests more often in Shona and Shona-English code-switched prompts than in English, suggesting AI safety protections built mainly for English …
- View project: Empirical Verification of Topological Phase Transitions During Grokking
Empirical Verification of Topological Phase Transitions During Grokking
Team grokking_code · Bengaluru
A multi-metric replication of the WeightWatcher framework analyzing the spectral properties and circuit complexity of neural networks during the grokking phase change. This project empirically isolates the exact mathematical inflection point where a network abandons dense memorization in favor of a sparse, …
- View project: PoliticAI
PoliticAI
Team Infinoid · Chennai
PoliticAI is a RAG-based political news and civic intelligence MVP designed to organize recent, source-linked political content into a safe and queryable knowledge system. It helps users retrieve grounded information about political developments, election-related reporting, misinformation trends, and region-specific …
- View project: Data safety for institutions
Data safety for institutions
Team OFFRUITY · Bindura
An overview of today’s existing institutions data safety strategy and how it can be transformed to protection within AI age
- View project: DDRR-Trust: Auditor de Gobernanza e Inteligencia Artificial para la Tokenización Inmobiliaria
DDRR-Trust: Auditor de Gobernanza e Inteligencia Artificial para la Tokenización Inmobiliaria
Team Llajwitas Tech · Santa Cruz - Bolivia
DDRR-Trust es una herramienta innovadora de Gobernanza de Inteligencia Artificial diseñada para proteger a los inversores en el emergente mercado de la tokenización inmobiliaria en Bolivia. El problema fundamental radica en que la transferencia de un activo digital en la blockchain no otorga derechos legales de …
- View project: Beyond English: Assessing the Robustness of LLM Safety Mechanisms Against Structural Jailbreaks in Spanish
Beyond English: Assessing the Robustness of LLM Safety Mechanisms Against Structural Jailbreaks in Spanish
Team Los Jailbreakers · Cali, Colombia
LLM safety guardrails are trained mainly to catch direct, single-turn harmful requests in English, leaving open whether structural attacks — long-context framing and multi-turn escalation — work just as well in Spanish, particularly in dialects like Bogotá Spanish. This matters for Latin America's emerging AI …
- View project: Writing for Combating Prompt Injection
Writing for Combating Prompt Injection
Team Bangalore_Electric_Sheep_Group · Bangalore
In this paper, we describe a method of protecting against prompt injection and preserving system prompt constraints from being over-ridden. We show that methods from dialog-state tracking and task-oriented dialog can mitigate the problems to a certain extent and result in significant improvements in robustness against …
- View project: AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts
AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts
Team AfriGuard AI · Nairobi, Kenya
Artificial intelligence is becoming part of how people learn, access information, make decisions, and participate in society. However, most AI safety testing is still designed around English conversations, leaving an important question unanswered: do AI systems remain reliable when people communicate in the …
- View project: REIA — Evaluador de Riesgo de Cumplimiento Legal en IA para Latinoamérica
REIA — Evaluador de Riesgo de Cumplimiento Legal en IA para Latinoamérica
Team CodByte19 · Santa cruz de la sierra/bolivia
REIA evalúa el riesgo de cumplimiento legal de un proyecto de IA en Latinoamérica. El usuario elige un país y sube la descripción de su proyecto (PDF, DOCX o texto); el sistema busca las normas vigentes de ese país —cada una con enlace a su fuente oficial—, analiza el proyecto contra ellas y devuelve: porcentaje de …
- View project: An Organization-Wide Safety Auditing Framework for Preventing and Detecting AI Sleeper Agents
An Organization-Wide Safety Auditing Framework for Preventing and Detecting AI Sleeper Agents
Team I worked solo · South Africa, Durban
I turned complex Anthropic AI Safety Research paper into a corporate auditing blueprint. It bridges the gap between human management and technical AI safety, making it a very strong research paper to probably secure myself a spot in your 6-8 weeks research studio with mentors. This topic was provided by you in the …
- View project: Developing a Context-Sensitive AI Governance Framework for Zambia
Developing a Context-Sensitive AI Governance Framework for Zambia
Team Greys team · Lusaka
Developing a Context-Sensitive AI Governance Framework for Zambia that facilitates the mitigation of gradual citizen disempowerment
- View project: Autonomous Institutions Safety Framework (AISF)
Autonomous Institutions Safety Framework (AISF)
Team from start to finish · Buenos Aires
Latin America may become one of the first regions where AI systems gain legally recognized authority over corporations, infrastructure, and capital management. Recent proposals in Argentina introducing Non-Human Corporations suggest a future where autonomous AI systems operate institutions traditionally managed by …
Overview
HACKATHON WINNERS
Congratulations to our winning teams, and thank you to everyone who submitted. We received 217 projects across three tracks, from hubs and participants across Latin America, Africa, and Asia, with support from Schmidt Sciences. The bar was high throughout. Each regional winner receives $1,000.
🌎 Latin America
🥇 Coldron by Leonardo Párraga, Angie Giraldo & Víctor Gelves
🥇 Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders by UAO SAFETY
🥇 Thought Anchors for Social Bias: Which Reasoning Steps Matter in Extended Thinking LLMs on Latin American Scenarios by Andres Felipe Mosquera Hernandez
🌏 Asia-Pacific
🥇 STEER: Mech interp-based white box attack on LLMs by Joshua Adrian Cahyono
🥇 Do Multilingual Vision-Language Models Abstain under Cross-Modal Conflict in Low-Resource Languages? by Anvesh Reddy Lankala, Vicky Feliren & Akansh Jain
🌍 Africa
🥇 Confidently Wrong: Measuring and Mitigating Calibration Risks in LLMs for African Languages by Team AIMS South Africa
🎖️ Honorable Mentions
Non-cash recognition for standout work, with these teams under consideration for the Apart Fellowship.
JusticIA
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
¿Por qué los agentes obedecen?
Local-Geometry Signals of Capability Emergence During Portuguese Grammar Acquisition
IndicViet-Safe
The Shapes of Bias in Spanish-Prompted LLMs and the Debiasing Prompt Scaffolds
Ufakazi
—————————————————————————————————————————————————
The Global South AI Safety Hackathon brings together researchers, engineers, and policy professionals across Latin America, Africa, and Asia to work on AI safety problems that matter most in their regions. Over one weekend, participants build tools, evaluations, and policy research addressing gaps the field has overlooked.
The hackathon is designed for and with the Global South. Participants compete within their region, not globally. The best teams from all regions are invited into the Apart Fellowship for continued research and mentorship.

Why the Global South?
AI safety research is concentrated in a handful of countries. None of the top 100 institutions by AI publication index, in either universities or companies, are based in Africa or Latin America (Chan et al., 2021). Meanwhile, AI risks hit differently in these regions: jailbreaks are more common in low-resource languages, and algorithmic bias trained on non-local data shows up in healthcare and hiring deployments.
This hackathon is not about bringing AI safety to the Global South. It is about bringing the Global South into AI safety. Researchers here have contextual knowledge the field needs: regulatory landscapes, language gaps, and institutional constraints that determine whether safety research actually works in practice.
Regional Tracks
Participants compete within their region. Each region has its own winners and prizes. Within each regional track, participants choose a sub-track (Technical AI Safety, AI Governance/Policy, or a locally-tailored sub-track defined by their hub). If you're based outside the three regions, you can still take part through the Open Track below.
Track 1: Latin America
Hubs in São Paulo, Buenos Aires, Bogotá, México: Merida-Guadalajara. Three winning teams ($3,000 total).
Brazil's AI bill (PL 2.338/2023), approved by the Senate and pending Chamber vote, includes a standalone human rights chapter that goes beyond the EU AI Act. Chile became the first country in the world to constitutionally protect neuro-rights. Colombia's CONPES 4144 established a national AI policy framework in early 2025.
We encourage submissions in three sub-tracks:
- Technical safety: AI fairness for Portuguese and Spanish language models, safety evaluations for AI systems deployed in Latin American contexts.
- Governance: regulatory analysis of emerging legislation across the region.
- Locally-tailored: problems defined by your hub that reflect local context.
- Open: any other project relevant to AI safety in the region that doesn't fit the categories above.
Track 2: Africa
Hub in Cape Town. One winning team ($1,000 total).
Africa faces risks from advanced AI that are not addressed through frontier AI safety efforts. African-context deployment-side risks include deepfake-driven electoral interference, data-colonial dependency on foreign infrastructure, compute and semiconductor scarcity constraining sovereign AI capacity, and large-scale labour market disruptions. This hackathon is an opportunity to prototype African-context solutions for misuse resilience, differential defence acceleration, and mitigating gradual disempowerment.
We encourage submissions in three sub-tracks:
- Policy: recommendations grounded in African regulatory and institutional contexts.
- Evals and benchmarks: work that builds on existing efforts and transfers to African contexts.
- Open: any other project relevant to AI safety in the region that doesn't fit the categories above.
Track 3: Asia
Hubs in Bengaluru (Electric Sheep), New Delhi (Secure AI Futures Lab), and Vietnam (Hanoi and Ho Chi Minh City). Two winning teams ($2,000 total).
Asia spans the full spectrum of AI governance approaches. India's "Seven Sutras" framework explicitly prioritizes innovation over restraint. China has enacted more sector-specific AI regulations than any other country. Vietnam's AI law took effect in March 2026, the first binding AI legislation in Southeast Asia. The ASEAN Guide on AI Governance and Ethics offers a voluntary regional framework. Projects can address cross-border governance harmonization, safety evaluations for non-English language models, technical AI safety research, or region-specific risk assessments.
We encourage submissions in three sub-tracks:
- Governance and geopolitics: AI compliance, AI sovereignty, national data security, compute diffusion and access.
- Socio-economic impacts: deepfakes and misinformation, caste bias, algorithmic bias across ethnic minorities, languages, and dialects, AI misdiagnosis in rural healthcare, gig platforms, labor displacement, gradual disempowerment, and cooperative AI.
- Technical safety: cybersecurity, multilingual, multimodal, and open-source models, small AI, and on-device AI.
- Open: any other project relevant to AI safety in the region that doesn't fit the categories above.
Track 4: Open
Based outside Latin America, Africa, and Asia? You can still take part. Submit your project to the Open Track and build the same way the regional tracks do: a tool, evaluation, or policy analysis on any AI safety problem that matters to you.
Who should participate?
- Founders and entrepreneurs
- AI safety researchers and engineers
- Machine learning researchers and engineers
- Policy researchers working on AI governance
- Software engineers interested in safety infrastructure
- Security researchers and red teamers
- Students and early-career researchers exploring AI safety
- Anyone working on AI and its impacts in the Global South
What you will do
Over three days, you will:
- Form teams and choose a regional track and sub-track
- Research and scope a specific problem using the provided resources
- Build a project: a tool, evaluation, policy analysis, or research contribution
- Submit a research report (PDF) documenting your approach, results, and implications
- Have your work reviewed by judges from AI safety organizations, universities, and policy institutions in your region
What happens next
After the hackathon, all submitted projects are reviewed by expert judges. Top projects receive prizes. The best teams may be invited into the Apart Fellowship for continued research and mentorship (subject to the asterisked caveat in Overview).
Organized by
Apart Research
Local hubs (8):
- Latin America: BAISH (Buenos Aires) | EA Brasil (São Paulo) | AI Safety Colombia (Bogotá) | AISMX (Mérida and Guadalajara)
- Africa: AI Safety South Africa (Cape Town)
- Asia: Electric Sheep (Bengaluru) | Secure AI Futures Lab (New Delhi) | AnToàn.AI (Hanoi and Ho Chi Minh City)
Supported by Schmidt Sciences.
Resources
Curated reading, tools, and project ideas. Start with Readings for All Tracks, then open your region. Each region lists suggested sub-tracks, reading, example projects, and local hubs. Shared tools and evaluation frameworks are at the bottom. You compete within your region; sub-tracks are guidance, not gates.
Project Ideas for all Tracks
Build with Adaption
Adaption builds adaptive AI infrastructure that lets researchers and builders train frontier models in days, not months. AutoScientist automates the full AI research and training loop, co-optimizing data and model recipes end-to-end until your model converges on your objective. Adaption supports 242 languages, making it one of the most linguistically inclusive AI platforms in the world. Participants can use AutoScientist directly to train and release their models and Adaptive Data to expand their datasets. 300 Adaption platform credits are included to get started.
Learn more about AutoScientist
Explore Adaptive Data
For questions you can join the Adaption Discord
Readings for All Tracks
AI Safety Foundations
- Apart Research: The Ultimate Guide to AI Safety Research Hackathons - How to approach a research hackathon and produce strong output.
- "Concrete Problems in AI Safety" (Amodei et al., 2016) - Foundational paper defining technical AI safety problems. Start here if new to the field.
- BlueDot Impact - Free AI safety courses covering alignment, governance, and technical safety.
AI Safety and the Global South
- "Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence" (Mohamed, Png, Isaac, 2020) - How colonial power dynamics shape AI development.
- "The Limits of Global Inclusion in AI Development" (Chan, Okolo, Terner, Wang, 2021) - None of the top 100 institutions at NeurIPS 2020 or ICML 2020 were from Africa or Latin America.
- "Why Global South Countries Need to Care About Highly Capable AI" (Abungu et al., CIGI, 2024) - The Global South must engage proactively with frontier AI risks.
- "Moving Beyond the Term 'Global South' in AI Ethics and Policy" (Stanford HAI) - Why the term flattens real differences in regional power and capacity.
- "Compute North vs. Compute South: The Uneven Possibilities of Compute-based AI Governance Around the Globe" (2024) - How compute access shapes who can govern AI.
AI Governance and Middle Powers
- "How AI Safety Is Getting Middle Powers Wrong" (Anton Leicht, 2025) - Argues for a pivot from frontier regulation to managing deployment risks.
- "The Early Death of International AI Governance" (Anton Leicht) - Where international AI governance is breaking down.
- "How Middle Powers Can Weather US and Chinese AI Dominance" (Chatham House, 2026) - Strategies for navigating US-China AI competition.
- "How Middle Powers May Prevent the Development of Artificial Superintelligence" (ControlAI) - A middle-power coalition theory of restraint.
- "The Race Worth Winning: Middle Powers in the Age of Machine Intelligence" (Leicht and Ball)
- "Strategic Choices for Middle Powers in Developing AI Capabilities: A Case Study of Singapore" (Cheng and Chong, Asian Security, 2025) - The 2C1D framework for middle-power AI strategy.
- "A Blueprint for Multinational Advanced AI Development" (Oxford Martin School)
- "Europe and the Geopolitics of AGI" (Centre for Future Generations)
Multilingual Safety and Evaluation (all regions)
- "Bridging the Multilingual Safety Divide" (Banerjee et al., 2026) - Safety alignment methods for Global South languages.
- "Code-Switching Red-Teaming" (Yoo, Yang, Lee, ACL 2025) - Mixed-language prompts achieve 46.7% more successful attacks. Tested across 10 languages.
- "Soteria: Language-Specific Safety Steering" (Banerjee et al., EMNLP 2025) - Lightweight per-language safety adjustment.
- "MrGuard: Multilingual Reasoning Guardrail" (Yang et al., EMNLP 2025) - Outperforms baselines by 15%+ across languages.
- "Refusal Direction Is Universal Across Languages" (Wang et al., NeurIPS 2025) - English refusal vectors transfer cross-lingually.
- "CREST: Cross-Lingual Safety Guardrails" (Bansal, Mishra, 2025) - A 0.5B safety classifier covering 100 languages, trained on only 13.
- "AraSafe: Arabic LLM Safety Benchmark" (EMNLP 2025) - 12K human-written plus 12K synthetic prompts; a replicable template.
- "DarkBench" (Kran et al., ICLR 2025 Oral) - Benchmark for dark patterns in LLMs.
Latin America
Hubs in Sao Paulo, Buenos Aires, Bogota, Merida, and Guadalajara. Three winning teams ($3,000 total). Brazil's AI bill (PL 2.338/2023) includes a standalone human rights chapter that goes beyond the EU AI Act. Chile became the first country to constitutionally protect neuro-rights. Colombia's CONPES 4144 established a national AI policy framework in early 2025.
Sub-track suggestions
- Technical AI safety: evaluations for agentic systems, mechanistic interpretability, AI fairness for Portuguese and Spanish language models.
- AI security: pipeline security (API, cloud), prompt injection and jailbreaks, AI control.
- Responsible AI: hallucination mitigation, behavioral audit, social impact evaluation.
- AI governance: AI policy and regulation analysis, system auditing and accountability, ecosystem monitoring, lethal autonomous weapons governance.
- Open: any other project relevant to AI safety in the region.
Reading
- Brazil AI Act (PL 2.338/2023) - Approved by the Senate, pending Chamber vote. Standalone human rights chapter.
- "SESGO: Spanish Evaluation of Stereotypical Generative Outputs" - 4,000+ prompts for Spanish and Latin American contexts.
Local hubs
- BAISH - Buenos Aires AI Safety Hub. Workshops, courses, scholarships.
- AISMX - AI Safety Mexico (Merida and Guadalajara).
- AI Safety Brazil - Brazil's AI safety community (Sao Paulo).
- Indice Latinoamericano de IA - Annual Latin American AI index covering infrastructure, R&D, and governance.
Africa
The global AI safety movement has focused on constraining frontier development in the US and China, leaving the deployment-side risks that affect Africa largely unaddressed. African countries face a distinctive risk profile: deepfake-driven electoral interference, data-colonial dependency on foreign infrastructure, compute and semiconductor scarcity, and large-scale labour market disruptions. Hub in Cape Town. One winning team ($1,000 total).
Sub-track suggestions
- Policy: comparative analysis of how African nations' draft AI policies address specific harms, with implementation-grounded recommendations.
- Evaluation and benchmarking: build on existing test suites or prototype new evals that transfer to African-context risks.
- Monitoring and tooling: prototypes for surfacing harms not currently captured, such as community-level monitoring.
- Open: any other project relevant to AI safety in the region.
Reading
Risk taxonomy and African AI safety agenda
- "Toward an African Agenda for AI Safety" (Segun et al., 2025) - Maps Africa's risk profile and proposes a five-point action plan, including an African AI Safety Institute.
- "Building Regional Capacity for AI Safety and Security in Africa" (Wiaterek, Abungu, Okolo, Brookings, 2025)
Disinformation and electoral integrity
- "African Democracy in the Era of Generative Disinformation" (Okolo, 2024) - Documented cases of AI propaganda in Nigeria, Burkina Faso, and Gabon.
- "Safeguarding African Democracies Against AI-Driven Disinformation" (CIPESA, 2025) - Civil-society analysis covering the African countries holding elections in 2025-2026.
- "Deepfake-Eval-2024" (Chandra et al., 2025) - Detectors drop 45-50% AUC on in-the-wild data; the collection method transfers to African electoral contexts.
- "XFacta" (2025) - Multimodal misinformation dataset from real fact-checking threads, all post-January 2024.
Policy and governance documents
- African Union Continental AI Strategy (2024) - The continental framework endorsed by the AU Executive Council; baseline for comparative African policy analysis.
- "The New Empire of AI: The Future of Global Inequality" (Rachel Adams, Polity Press, 2024) - How AI reshapes power between states, corporations, and Global South communities.
African-language NLP benchmarks
- "AfriSenti" (Muhammad et al., 2023) - 110,000+ tweets across 14 African languages annotated for sentiment.
- "IrokoBench" (Adelani et al., 2024) - Human-translated benchmark for 17 low-resource African languages (includes AfriMMLU).
- "AfroBench" (Ahuja et al., 2023) - 15 tasks, 22 datasets, 64 languages. The broadest map of what exists and what is missing.
AI safety evaluation for African languages
- "UbuntuGuard: Culturally-Grounded Safety Benchmark for African Languages" (Abdullahi et al., 2026) - First African-language safety benchmark, built by 155 domain experts across 10 languages.
- "LSR: Linguistic Safety Robustness Benchmark" (2026) - Cross-lingual refusal degradation for Yoruba, Hausa, Igbo, and Igala; built on the UK AISI Inspect framework.
Red-teaming and participatory harm monitoring
- "HarmBench" (Mazeika et al., 2024) - Standardized red-teaming harness across 33 LLMs; load African-language prompts or local harm categories.
- "Red Teaming Language Models to Reduce Harms" (Ganguli et al., Anthropic, 2022) - 38,961 human-written attacks, open dataset, extensible harm taxonomy.
- "STAR: SocioTechnical Approach to Red Teaming" (Vidgen et al., 2024) - Demographic matching between red-teamers and the harms being probed.
- "Ask What Your Country Can Do For You" (Kenway et al., 2025) - Portable methodology for multilingual, multicultural public red-teaming.
- "Particip-AI" (Mun et al., 2024) - Four-step framework for engaging lay communities in AI harm anticipation; tested with 295 participants.
- "Participatory Red-Teaming with Targets of Stereotyping" (Kim et al., 2026) - Engaging affected communities as red-teamers, and the costs of doing so.
Broader evaluation methodology
- "Holistic Evaluation of Language Models (HELM)" (Liang et al., 2022) - Multi-scenario, multi-metric benchmarking template that makes coverage gaps visible.
- "Constitutional AI: Harmlessness from AI Feedback" (Bai et al., 2022) - The constitution can be rewritten to embed locally-derived African norms.
- "A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms" (Abercrombie et al., 2024) - Harm taxonomy built from 39 real incidents, legible to NGOs and civil society.
- "CulturalBench" (Chiu et al., 2024) - 1,696 cultural-knowledge questions across 45 regions; pipeline adaptable for Africa-specific questions.
Institutions Doing Ongoing Work
- UCT African Hub on AI Safety, Peace and Security
- Global Centre on AI Governance
- Tech Governance Project
- CIPESA
- The Collective Intelligence Project, especially their weval work.
Resources
- AI Safety South Africa - capacity building, co-working hub and research\
- African Hub on AI Safety, Peace and Security - research, policy development, and capacity building
- ILINA Program - research, policy engagement and talent development
Asia
The world's most dynamic region is also its most exposed. Asia holds half the world's people, most of its informal workers, and its sharpest geopolitical fault lines. Asia spans the full spectrum of AI governance approaches. India's "Seven Sutras" framework prioritizes innovation over restraint. Vietnam's AI law took effect in March 2026, the first binding AI legislation in Southeast Asia. The ASEAN Guide on AI Governance and Ethics offers a voluntary regional framework. Hubs in Bengaluru (Electric Sheep), New Delhi (Secure AI Futures Lab), and Vietnam (Hanoi and Ho Chi Minh City). Two winning teams ($2,000 total)
Sub-track suggestions
- Governance and geopolitics: AI compliance, AI sovereignty, national data security, compute diffusion and access.
- Socio-economic impacts: deepfakes and misinformation, caste bias, algorithmic bias across ethnic minorities, languages, and dialects, AI misdiagnosis in rural healthcare, gig platforms, labor displacement, gradual disempowerment, and cooperative AI.
- Technical safety: cybersecurity, multilingual, multimodal, and open-source models, small AI, and edge AI.
- Open: any other project relevant to AI safety in the region.
Reading
Governance and geopolitics: AI sovereignty, compute, and access
- "India AI Governance Guidelines" (the "Seven Sutras" framework)
- "China AI Safety Governance Framework v2.0" (TC260, 2025)
- "ASEAN Guide on AI Governance and Ethics" (2024)
- "Is AI Sovereignty Possible? Balancing Autonomy and Interdependence" (Brookings, 2026)
- "AI Compute Sovereignty: Infrastructure Control Across Territories, Cloud Providers, and Accelerators" (Hawkins et al.)
- "The Geopolitics of AI and the Rise of Digital Sovereignty" (Brookings, 2023)
- "From Sovereignty to Coordination: Rethinking National AI Strategy" (The Economy Research, 2026)
Socio-economic impacts: labor, gradual disempowerment, cooperative AI
- "Artificial Intelligence's Creation and Displacement of Labor Demand" (Technological Forecasting and Social Change, 2024)
- "The Future of Work, Artificial Intelligence, and Digital Government" (Asian Development Bank Institute)
- "Digital Progress and Trends Report 2025: Strengthening AI Foundations" (World Bank)
- "Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development" (Kulveit et al., 2025)
Technical safety: cybersecurity, multimodal and open models, small and edge AI
- "Multimodal AI: A Guide to Open-Source Vision-Language Models" (BentoML)
- "Risks and Opportunities of Open-Source Generative AI"
- "OWASP Top 10 for LLM Applications"
- "Beyond the Tip of Efficiency: Submerged Threats of Jailbreak Attacks in Small Language Models" (ACL 2025)
- "Small but Dangerous: Evaluating and Mitigating Jailbreak Vulnerabilities in Small Language Models"
- "Small Models, Big Problems: Why Your AI Agents Might Be Sitting Ducks" (Enkrypt AI)
- "Edge AI: A Taxonomy, Systematic Review and Future Directions"
- "Security Risks of AI Hardware for Personal and Edge Computing Devices" (NCC Group)
- "Privacy and Security Vulnerabilities in Edge Intelligence"
- "IndoSafety" (EMNLP 2025) - Culturally grounded Indonesian safety benchmark, replicable for Vietnamese.
- "ThaiSafetyBench" - Thai-culture-specific safety taxonomy, a template for Southeast Asian benchmarks.
- "DECASTE: Caste Stereotypes in LLMs" (IJCAI 2025) - First systematic benchmark for caste bias in AI.
Asia as the most dynamic and most exposed region
- "The Indo-Pacific as the Epicenter of AI Risk" (NBR)
- "AI Safety Needs Southeast Asia's Expertise and Engagement" (Brookings)
- "Global Majority Is Building International AI Governance Through Cooperation, Not Competition" (Bennett Institute, ASEAN focus)
- "AI Governance in East Asia" (in Inter-Asian Law, Cambridge University Press)
India-specific: deepfakes, caste bias, rural health, and the AI Safety Institute
- "Deepfakes: How India Is Tackling Misinformation During Elections" (World Economic Forum, 2024)
- "India's Generative AI Election Pilot Shows AI in Campaigns Is Here to Stay" (Center for Media Engagement, 2024)
- "MeitY-UNESCO Stakeholder Consultation on AI Safety and Ethics" (IndiaAI, 2024)
Vietnam-specific: the Law on AI and national strategy
- "National Strategy for AI Research, Development and Application through 2030" (2021, English)
- "Law No. 134/2025/QH15: Law on Artificial Intelligence" (December 2025)
- "Decree No. 142/2026/ND-CP: Implementing the Law on Artificial Intelligence" (April 2026)
- "Vietnam National AI Ethics Framework" (2025)
- "Strategy for the Development of Vietnam's Semiconductor Industry to 2030, Vision 2050" (2024, English summary)
Local hubs
- Electric Sheep - Bengaluru-based AI safety hub.
- AI Safety India - Programs, research, network.
- Antoàn.ai - AI Safety Vietnam
- Secure AI Futures Lab - AI Safety Network and Research Lab.
- AI Safety Asia - Pan-Asian AI safety network.
Shared Tools and Evaluation Frameworks
Safety evaluation
- DeepSafe Toolkit (AI45Lab) - Integrates 20+ safety benchmarks with a YAML-driven pipeline.
- Promptfoo - Open-source LLM evaluation with built-in multilingual jailbreak testing.
- LightEval (Hugging Face) - Lightweight, configurable LLM evaluation toolkit.
- PolyGuard (COLM 2025) - Multilingual safety classifier for 17 languages.
- SALAD-Bench (Shanghai AI Lab) - Hierarchical safety benchmark with native Chinese-language coverage.
- ControlArena (UK AISI / Redwood) - AI control evaluation framework for agent safety scenarios.
Governance trackers and bias auditing
- OECD.AI Policy Navigator - AI policy initiatives from 80+ jurisdictions.
- UNESCO Global AI Ethics Observatory - Country profiles for 40+ nations.
- UNESCO Ethical Impact Assessment - Tool for evaluating AI systems against UNESCO's AI Ethics Recommendation.
- AGILE Index 2025 - AI governance scores for 40 countries.
- Oxford Insights Government AI Readiness Index - 195 governments scored on AI readiness.
- Global Index on Responsible AI - Responsible AI commitments in 138 countries.
- Africa AI Policy Tool - 1,465+ national AI strategies and policies across Africa, filterable by country.
- African AI and Equality Toolbox (Alan Turing Institute and OHCHR) - Human rights-based AI lifecycle framework adapted for Africa.
- Aequitas - Open-source Python toolkit for auditing ML models for discrimination.
- Algorithm Audit: Unsupervised Bias Detection - Detects bias without protected-attribute data.
Guidelines
Judging Criteria
Dimension 1: Impact Potential & Innovation
How much would this matter for AI safety if it worked? How innovative is it?
For scores of 4-5: is this actually new to the field, or replicating recent work?
| Score | Description |
|---|---|
| 1 | Negligible. No clear problem addressed, or no meaningful novelty. |
| 2 | Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best. |
| 3 | Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools. |
| 4 | Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on. |
| 5 | Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area. |
Dimension 2: Execution Quality
How sound are methodology, implementation, and findings?
| Score | Description |
|---|---|
| 1 | Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work. |
| 2 | Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation. |
| 3 | Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions. |
| 4 | Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work. |
| 5 | Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation. |
Dimension 3: Presentation & Clarity
How clearly are work, findings, and impact potential communicated?
| Score | Description |
|---|---|
| 1 | Incomprehensible. Cannot determine what the project is actually claiming or doing. |
| 2 | Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points. |
| 3 | Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations. |
| 4 | Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly. |
| 5 | Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work. |
Submission Requirements
A complete submission includes:
- A research report in PDF format using the official submission template
- A project title and brief abstract (150 words max)
- Author names and affiliations for all team members
- An indication of which regional track and sub-track your project addresses
Important: Include a section called "Limitations and Dual-Use Considerations" that addresses limitations of your approach, potential misuse risks, and suggestions for future improvements.
Report structure (recommended):
There is no hard page limit. Most winning projects are 4-8 pages. The report should include:
- Introduction: What problem did you address? Why does it matter for AI safety in your region or globally?
- Related Work: What existing work does your project build on?
- Methodology: What did you build or test? Describe your approach in enough detail for someone to replicate it
- Results: What did you find? Include quantitative results where possible
- Discussion: What are the implications? What are the limitations? What would you do with more time?
- References: Cite relevant prior work
- Do not submit the whole report in Italic
Important notes:
- You can submit as an individual or as a team
- You can build on existing work, but you MUST clearly identify what is NEW work done during the hackathon
- Submit through the official submission form on the hackathon page
- If you run into submission issues, DM Kamil on Discord or email sprints@apartresearch.com
- Have team names, emails, Discord handles, project title, abstract, and PDF ready before starting the form.
- Reports can be in English, Spanish, or Portuguese. English is preferred if your team is comfortable with it, but we will not penalize non-English submissions.
- Due to the high volume of submissions, we may not be able to provide written feedback for every project, but all projects will be evaluated.
Submission Template => Link
Discord Server Invite => Link
Frequently Asked Questions
Eligibility and Format
Q: Do I have to be from the Global South to participate?
We use Britannica's definition of the Global South. If you're based in Latin America, Africa, the Middle East, or Asia/Oceania (excluding Japan, South Korea, Taiwan, Australia, New Zealand), you're in scope.
If you're based outside these regions, you're still welcome: submit your project to the Open Track (Track 4). If you're not sure where you fit, sign up and we'll sort it out.
Q: Do I need a research background or coding experience to participate?
No formal requirement. Past sprints have had grad students, undergrads, industry engineers, and self-taught researchers all submit winning projects. What helps: ability to scope a small experiment, write up findings clearly in a short PDF. We provide resources, mentor support during the weekend, and the rubric publicly so you know what's being evaluated. If you're unsure whether you'd fit, ask in the Frequently Asked Questions
Answers below are from the Apart master FAQ database (Kamil-approved). For anything not covered, ask on Discord help-desk or email sprints@apartresearch.com.
Eligibility and Format
Q: Do I have to be from the Global South to participate?
We use Britannica's definition of the Global South. If you're based in Latin America, Africa, the Middle East, or Asia/Oceania (excluding Japan, South Korea, Taiwan, Australia, New Zealand), you're in scope. If you're not sure, sign up and we'll sort it out.
Q: Do I need a research background or coding experience to participate?
No formal requirement. Past sprints have had grad students, undergrads, industry engineers, and self-taught researchers all submit winning projects. What helps: ability to scope a small experiment, write up findings clearly in a short PDF. We provide resources, mentor support during the weekend, and the rubric publicly so you know what's being evaluated. If you're unsure whether you'd fit, ask in the Discord help-desk.
Q: Is the hackathon remote or in-person? Can I attend remotely if there's a local hub in my city?
Default is remote. We also have local in-person hubs where people optionally gather for the weekend. You can join remotely from anywhere even if there's a hub in your city. Hubs are optional and additive, not a separate event. Talks are livestreamed to remote participants; judging is always asynchronous and online.
Q: When is the hackathon?
June 19-21, 2026. It starts Friday evening. Submissions close Sunday June 21 at 11:59 PM AoE (Anywhere on Earth).
Teams and Tracks
Q: Can I participate solo, or is it teams only? What's the max team size?
Both solo and team submissions are welcome. Maximum team size is 5. Smaller teams (1-3) are common and have won prizes. If you're solo and want a team, post in the Discord #team-formation channel a week or two before the event. We don't auto-match teams; it's organic.
Q: How strict are you on the tracks? Can my project fit between two or none of them?
Tracks are guidance, not gates. If your project genuinely lives between two tracks, pick the closest one and explain the cross-track angle in your submission. Judges score on Impact, Execution, and Presentation across all tracks; nobody gets disqualified for fit. If it really fits none of the tracks but is in scope for the hackathon's overall topic, submit anyway and Kamil will route it. Better to submit and have us re-bin than to skip.
Q: Can a team have members from different regions?
Yes. Choose the regional track that best fits your project. Your team can include members from any location. You compete for prizes in whichever regional track you select.
Submissions
Q: What do I submit?
A research report in PDF format using the submission template. Think of it as a mini research paper documenting your problem, approach, results, and implications. Not a product demo.
Q: Can I build on existing work, or does everything have to be from scratch during the hackathon?
Yes, you can build on existing work, as long as the contribution you make during the hackathon is clearly identifiable in your submission. Judges look for what's new this round. Cite prior work in your PDF (anywhere from a sentence to a methods paragraph; whatever makes the delta clear). The Project Eligibility note in our rubric covers this. Submissions that don't show a clear during-event delta tend to score lower on Execution.
Q: Do you recommend a particular referencing/citation style for the project write-up?
No required style. Pick whichever you are comfortable with (APA, MLA, Chicago, IEEE, etc.) and stay consistent throughout the paper. The judges care about the substance of the work, not the citation format.
Q: Can I submit an unfinished project?
Yes. Submitting something unfinished is always better than not submitting. Judges evaluate what you accomplished during the hackathon timeframe.
Judging, Results, and Prizes
Q: When are winners announced after the hackathon ends?
Roughly 1-3 weeks after the event. The timeline: judges receive assignments Monday or Tuesday after the weekend, reviews are due the following Sunday, then we tally and rank, then we send winner notifications. For biosecurity-related events we add a hazard review step before announcing publicly, which can extend the timeline. We send everyone reviewer feedback after winners are notified.
Q: How are prizes awarded?
Prizes are awarded per region, not globally. Latin America has 3 winning teams, Asia has 2, and Africa has 1. You compete with participants in your regional track.
Q: How and when do prize payments arrive after the hackathon?
Winners are announced 1-3 weeks after the event. After announcement, your team agrees on a prize split and nominates one person to receive the funds. That nominee fills out a payment form and is onboarded to our payment processor (Ashgro/Ramp). Payment lands roughly 2-4 weeks after the form is submitted, depending on the payout method. The nominee handles distribution to the rest of the team. International transfers may take longer.
Support
Q: How do I get help during the hackathon?
Three ways:
- Discord help-desk channel (tag @Kamil Alaa)
- DM Kamil on Discord
- Email: sprints@apartresearch.com
Schedule
| Wednesday June 17 | ||
|---|---|---|
| 23:00 UTC | Miguel A. Peñaloza, Statistician, LLM evaluation researcher | Recording |
| Thursday June 18 | ||
| 02:00 UTC | Valerie Pang, Program Manager, Singapore AI Safety Hub | Recording |
| 11:15 UTC | Tianyi Qiu, Alignment researcher, Oxford HAI / Stanford CS | Recording |
| 13:30 UTC | Akash Kundu, Cooperative AI Research Fellow | Recording |
| 14:30 UTC | Carlos Giudice, Research Engineer, EquiStamp | Recording |
| 16:00 UTC | James Fox, Senior Science Associate, Schmidt Sciences | Recording |
| 17:00 UTC | Srinivas Pochincharla, Senior Technical Program Manager, Amazon | Recording |
| 18:00 UTC | Samuel Segun, CEO, Quelox Labs | Recording |
| Friday June 19 | ||
| 12:00 UTC | Clement Neo, Founder, Neo Research | Recording |
| 13:00 UTC | Hieu Vu, language-model interpretability researcher | Recording |
| 15:00 UTC | Naveen Kumar Gond, BASIS (UC Berkeley) | Recording |
| 16:30 UTC | Sang Truong, PhD candidate, Stanford AI Lab | Recording |
| 17:30 UTC | Roshni Lulla, Co-founder & CRO, Institute for Humane Robotics | Recording |
| 18:30 UTC | Jasmine Li, Research Scientist, SaferAI | Recording |
| 20:00 UTC | Pradyumna Shyama Prasad, Evaluations Fellow, Elicit | Recording |
| Saturday June 20 | ||
| 00:00 UTC | Sissi de la Peña, Founder, The Dot Network | Recording |
| 01:00 UTC | Tan Zhi Xuan, Presidential Young Professor, NUS Computer Science | Recording |
| 14:00 UTC | Janhavi Khindkar, Bhashini; and leads ValueShift Research | Recording |
| 17:00 UTC | Hemanth Badabagni, Senior Forward Deploy Engineer at Databricks | Recording |
| 18:00 UTC | Juan Roberto Hernández Villalobos, AI Consultant, UNESCO | Recording |
| Sunday June 21 | ||
| 01:00 UTC | Amaia Amezaga, Founder, Sattva Labs | Recording |
| 18:00 UTC | Isabella Rodrigues Martins, Cybersecurity professional | Recording |
| 20:00 UTC | Juan Felipe Cerón Uribe, AI Alignment Research Engineer, OpenAI | Recording |
| Logistics Presentation + QA | Recording | |
| Submission Deadline is Sunday End of Day Anywhere on Earth |
More Talks are being scheduled
Speakers

James Fox
Speaker
James Fox is a Senior Science Associate at Schmidt Sciences, where he leads the Science of Trustworthy AI program and supports the broader mission of the AI and Advanced Computing Institute. Prior to joining Schmidt Sciences, James served as Research Director of the London Initiative for Safe AI (LISA). He earned his DPhil in Computer Science from the University of Oxford, supervised by Tom Everitt, Michael Wooldridge & Alessandro Abate; research on game theory, causality, and reinforcement learning. MSci/BA Natural Sciences (Physics), Cambridge.

Juan Felipe Cerón Uribe
Speaker and Judge
Juan Felipe Cerón Uribe is an AI Alignment Research Engineer at OpenAI in San Francisco, where he works on mitigating existential risks from artificial intelligence. He previously built AI systems at Factored and was a research intern at the Stanford Existential Risks Initiative.

Tan Zhi Xuan
Speaker
Tan Zhi Xuan is a Presidential Young Professor in the NUS Department of Computer Science. Xuan's research focuses on scaling cooperative intelligence via rational and model-based AI, spanning probabilistic programming, model-based planning, Bayesian inference, AI alignment, and computational cognitive science, and leads the Cooperative Systems & Intelligence (CoSI) lab.

Samuel Segun
Speaker and Judge
Dr. Samuel T. Segun is an expert in AI safety, ethics and technical governance, advising governments, multilateral institutions, and research networks on responsible AI. He is an Honorary Senior Lecturer and a Senior Research member of the African Hub on AI Safety, Peace and Security at the University of Cape Town, Editor-in-Chief of The Algorithmic Review, and sits on the Board of Advisors for VerifyWise. Previously he was Senior Consultant on AI Safety to the UN Office of Counter-Terrorism (UNOCT) and UNICRI. He is the author/editor of several articles and books (Palgrave, Nature, Springer, Cambridge), including Selected Issues in the Ethics of AI (2022) and AI, Ethics and Policy Governance in Africa (2024).

Umut Pajaro Velasquez
Speaker and Judge
Umut Pajaro Velasquez is President of the Internet Society Colombia Chapter and holds an MA in Cultural Studies, researching education, digital rights, and ethical AI design and development, with a focus on mitigating bias against marginalized groups. Umut has chaired Internet Governance groups, including the Internet Society Gender Standing Group.

Sang Truong
Speaker
Sang Truong develops foundations for AI Measurement Science, drawing on probabilistic machine learning, measurement theory, and mechanism design to improve how we evaluate AI systems. His work supports the development of AI systems that serve people across diverse backgrounds and needs. He is a PhD candidate in Computer Science at the Stanford AI Lab with Sanmi Koyejo and Nick Haber.
Show 19 moreShow fewer

Tianyi Qiu
Speaker and Judge
Tianyi is an alignment researcher at Oxford HAI / Stanford CS and former Anthropic AI Safety Fellow. He develops human-in-the-loop training interventions that make AI foster, not undermine, growth in human ideas. Research projects he led have received two Best Paper Awards (ACL'25; NeurIPS'24 Pluralistic Alignment Workshop).

Akash Kundu
Speaker
Akash Kundu is an AI safety researcher from Kolkata with publications at NeurIPS and ICLR. He has worked with organisations including FAR.AI and Apart Research (as a lab fellow) on AI safety and LLM evaluation, and is currently in the extension phase of the Cooperative AI Research Fellowship.

Amaia Amézaga
Speaker and Judge
Amaia Amézaga is an independent researcher and the founder of Sattva Labs, an initiative dedicated to AI safety, model evaluation, and human-AI interaction. Her approach combines interaction design (HCI), multicultural and multilingual communication, and qualitative analysis to explore how people interpret, use, and relate to artificial intelligence systems.

Carlos Giudice
Speaker
Carlos Giudice is a research engineer at EquiStamp, contracting for Redwood Research on AI control, where he helped build LinuxArena. Before AI safety he spent six years as an ML engineer at MercadoLibre (Latin America's largest e-commerce company) and a year at CERN simulating data-transfer networks. He co-organizes BAISH, the Buenos Aires AI Safety Hub.

Hieu Vu
Speaker and Judge
An engineer turned independent researcher studying the interpretability of language models.

Luis Enrique Urtubey De Césaris
Speaker and Judge
Luis Enrique Urtubey De Césaris is Director of Strategy at CEGIA, the Centro de Estudos em Governança de Inteligência Artificial, where he leads work on AI governance and field-building. His current focus areas include PL 2338/2023, Brazil's framework AI bill; the creation of a Brazilian AI Safety Institute; international AI governance coordination; and capacity-building in AI governance and technical AI safety.

Miguel A. Peñaloza
Speaker
Miguel A. Peñaloza is a professional statistician whose work focuses on developing benchmarks to assess the capabilities of Large Language Models (LLMs). Currently, he is designing methodologies aimed at making AI system evaluations more inclusive, representative, and globally accessible.

Naveen Kumar Gond
Speaker
Naveen Kumar Gond is an alumnus of the BASIS AI Safety Policy Fellowship at UC Berkeley with a background in public policy and governance. His focus is on global AI governance and frontier AI risk, particularly how international institutions and national regulators can develop coherent frameworks for advanced AI systems. He is currently building his research and policy practice at the intersection of AI alignment, responsible AI development, and regulatory affairs.

Roshni Lulla
Speaker
Roshni is co-founder and CRO of the Institute for Humane Robotics (IHR), working on guidelines for human-robot interactions, and a neuroscientist (PhD, Brain & Cognitive Sciences, USC, under Antonio Damasio and Jonas Kaplan). Bridges affective neuroscience, moral cognition, and AI safety.

Sissi de la Peña
Speaker
Sissi de la Peña is the Director of The Dot Network and served as a 2024-2025 Tech Policy Fellow at UC Berkeley. She has spent over 20 years in Latin American digital and AI policy, from USMCA digital-trade negotiations to representing major technology companies in regional regulatory forums, and was named Responsible AI Leader of the Year at the 2025 Women in AI Awards.

Srinivas Pochincharla
Speaker and Judge
Srinivas Pochincharla is an engineering leader and Senior Technical Program Manager at Amazon with over 12 years of experience in software engineering, enterprise data architecture, and AI/ML-powered product development. He specializes in leading complex, cross-functional initiatives at the intersection of large-scale data platforms and intelligent automation, with a particular focus on Generative AI and enterprise data transformation.

Valerie Pang
Speaker
Valerie is the Program Manager at Singapore AI Safety Hub where she handles memberships, marketing, events and operations. She has 4 years of experience working in tech companies in the US and Singapore, and 2 years of international non-profit experience working in corporate engagement. She has experience organising events and conferences with hundreds of attendees.

Janhavi Khindkar
Speaker, Judge and Mentor
Janhavi Khindkar is an Applied AI Researcher and Engineer working on Bhashini, India's national multilingual AI platform under MeitY, where she works on model optimization, fine-tuning, and deployment for low-resource Indic languages at scale. She also leads ValueShift Research, an independent AI safety collaboration focused on mechanistic interpretability and AI control. Her work sits at the intersection of applied ML infrastructure and AI safety, with a particular interest in how safety alignment behaves across languages and cultural contexts.

Hemanth Badabagni
Speaker and Judge
Hemanth Badabagni is a technology leader specializing in data engineering, analytics, and AI-enabled data platforms. His work focuses on data governance, semantic systems, and enterprise AI, with an emphasis on building trusted, scalable foundations for data-driven decision-making.

Isabela Rodrigues Martins
Speaker
Isabela is a cybersecurity professional with four years of experience focused on protecting data and systems. She continues to build her skills through specialized training, including AWS, the Microsoft SC-900 certification (via Womcy LATAM in partnership with Microsoft), and Harvard's CS50 Introduction to Computer Science (offered in Brazil through Fundacao Estudar), which covered JavaScript, Python, CSS, and HTML.
Judges and mentors

Aniruddh Bhaduri
Mentor

Isabella Luong
Mentor and Judge
- (opens in new tab)

Trang Pham
Mentor and Judge
- (opens in new tab)

Neeraj Kumar Singh Beshane
Mentor and Judge
- (opens in new tab)

Haakon Huynh
Mentor and Judge
- (opens in new tab)

Katerine Hernandez
Mentor and Judge
- (opens in new tab)

Nguyen Duy Tung
Mentor and Judge
- (opens in new tab)

David Williams-King
Mentor and Judge
- (opens in new tab)

Claude Formanek
Mentor and Judge
- (opens in new tab)

Aditya Arpitha Prasad
Mentor and Judge
- (opens in new tab)

Anusha Mujumdar
Mentor and Judge
- (opens in new tab)

Amol Walvekar
Judge
Show 76 moreShow fewer

Aditya Thakur
Judge

Vivek Kotecha
Judge
- (opens in new tab)

Hardik Chawla
Judge

Agni Tripathi
Judge

sushant awasthi
Judge

Nihal K
Judge
- (opens in new tab)

David Salinas
Judge
- (opens in new tab)

Abhishek Das (Salesforce)
Judge
- (opens in new tab)

Varun J. Vincent
Judge

Temi Oloyede
Judge

Saraswati Mishra
Judge
- (opens in new tab)

Nikhil Reddy Pallepati
Judge

Wairagala Wakabi
Judge

Jady Pamella
Judge

Syed Anas Mohiuddin
Judge
- (opens in new tab)

Soumya Jain
Judge

Vanshika Gupta
Judge
- (opens in new tab)

Luis Cosio
Judge
- (opens in new tab)

Nikita Lokhmachev
Judge
- (opens in new tab)

Cibeles Garcia Burt
Judge
- (opens in new tab)

Vincent Mai
Judge
- (opens in new tab)

Camilla Balbis
Judge
- (opens in new tab)

Melissa Robles
Judge
- (opens in new tab)

Catalina Bernal
Judge

Juan Lievano-Karim
Judge
- (opens in new tab)

Steve Hege
Judge
- (opens in new tab)

Wanda Muñoz
Judge
- (opens in new tab)

Alejandro Acelas
Judge
- (opens in new tab)

Monica Ulloa
Judge

Jyoti Lakra
Judge

Suneet Malhotra
Judge

Vashishtha Patil
Judge
- (opens in new tab)

Ashwin Pai
Judge

Gaurav Kumar Sinha
Judge

San Krish
Judge

Syam Dondapati
Judge
- (opens in new tab)

Tzu Kit Chan
Judge and Mentor
- (opens in new tab)

Yogesh Thanvi
Judge

Ankit Arya
Judge

Ksheeraj Sai Vepuri
Judge

Ashwin Krishnappa Kumar
Judge

Naga Sujitha Vummaneni
Judge

Rahul Nambiar
Judge

Deepesh Khanna
Judge

Shantanu Bhatt
Judge
- (opens in new tab)

Surbhi Madan
Judge

Vivek Kumar
Judge

Juan Djuwadi
Judge
- (opens in new tab)

Vasilii Kondyrev
Judge

Pavel Sikachev
Judge

Joanna Wiaterek
Judge

Mark Gaffley
Judge

Jesús M. Siqueiros
Judge

Amanda Isbosseth Guzmán Aceves
Judge

Marta Kosmyna
Judge

Osmani Redondo
Judge

Fernando Castillo
Judge

Marcos Galván López
Judge

Silvia Fernández Sabido
Judge

Ivete Sánchez Bravo
Judge
- (opens in new tab)

Germán López-Ardila
Judge
- (opens in new tab)

Diego Ortiz Barbosa
Judge
- (opens in new tab)

Diego Gomez
Judge
- (opens in new tab)

Maria Paula Mujica
Judge

Rajanikant Vellaturi
Judge

Phani Harish Wajjala
Judge
- (opens in new tab)

Linh Nguyen
Judge
- (opens in new tab)

Ajay Devineni
Judge

Venkata Sangaraju
Judge

David Mazumdar
Judge
- (opens in new tab)

Jonathan Ng
Judge

Royston Monteiro
Judge

Suprita Shankar
Judge

Anchit Jhingan
Judge

Pratham Patkar
Judge

Michelle Malonza
Judge
Organizers
- (opens in new tab)

Kamil Alaa
Organizer

Ivan Martucci Franco
Hub Organizer EA Brazil
- (opens in new tab)

Tegan Green
Hub Organizer AIS South Africa

Nguyen Tran
Hub Organizer AnToàn.AI and Judge

Jose Gelves
Hub Organizer AIS Colombia

Max Pinelo
Hub Organizer AIS Mexico and Mentor

Isabel Camara
Hub Organizer AIS Mexico

Janeth Valdivia
Hub Organizer AIS Mexico

Angel Tenorio
Hub Organizer AIS Mexico and Mentor

Marco Guzman
Hub Organizer AIS Mexico

Dexter Gomez
Hub Organizer AIS Mexico

Basil Labib
Hub Organizer SAFL
Show 2 moreShow fewer

Ash Singh
Hub Organizer Electric Sheep

Diksha Singh
Hub Organizer Electric Sheep
Local sites
Apart Global South AI Safety Hackathon — México
A collaborative space connecting students, researchers, engineers, policymakers, industry professionals, and academic communities across Mexico with the global AI safety ecosystem. Participants will collaborate on technical and governance projects addressing the challenges and opportunities of AI safety in the Global South. Mexico will host two in-person hubs in Mérida and Guadalajara, while participants from across the country are also welcome to join remotely. Join us!
Event page: Apart Global South AI Safety Hackathon — México (opens in new tab)Apart Global South AI Safety Hackathon — Santa Cruz Hub
Local hub for the Apart Research Global South AI Safety Hackathon (June 19–21, 2026). A collaborative workspace connecting Bolivian researchers, engineers, and students with the global AI safety community. We will stream Apart's keynotes and mentorship sessions while working on technical and governance projects in teams.
Event page: Apart Global South AI Safety Hackathon — Santa Cruz Hub (opens in new tab)Global South AI Safety Hackathon — Bogotá Hub
Local Bogotá hub for the Global South AI Safety Hackathon, coordinated by AI Safety Colombia. Participants will work in teams on technical AI safety, AI security, and responsible AI/governance projects connected to Latin American contexts, with access to talks, mentorship, and the broader Apart Research sprint.
Event page: Global South AI Safety Hackathon — Bogotá Hub (opens in new tab)Global South AI Safety Hackathon: Apart x AI Safety Zimbabwe
The Global South AI Safety Hackathon (African track) brings together researchers, students and builders in Harare to develop evaluations, policy and tools for African AI risk contexts, hosted by AI Safety Zimbabwe, EA Zimbabwe and the AI Collective. Exact venue shared with registered participants.
Event page: Global South AI Safety Hackathon: Apart x AI Safety Zimbabwe (opens in new tab)Global South AI Safety Hackathon: Buenos Aires
Sumate al nodo en Buenos Aires de la hackathon del sur global, organizado por BAISH (https://baish.com.ar/).
Event page: Global South AI Safety Hackathon: Buenos Aires (opens in new tab)Global South AI Safety Hackathon: Cape Town
Develop context appropriate solutions to policy and evaluations for the African AI risk response, focussing on policy or technical interventions against the AI harms most likely to manifest in African deployment contexts.
Event page: Global South AI Safety Hackathon: Cape Town (opens in new tab)Global South AI Safety Hackathon: Da Nang
AnToàn.AI (AI Safety Vietnam) is hosting the Da Nang jam site for the Global South Al Safety Hackathon. We will provide coworking spaces & refreshments on-site on June 20 and 21 for participants, along with streaming Apart's keynote and mentorship sessions.
Event page: Global South AI Safety Hackathon: Da Nang (opens in new tab)Global South AI Safety Hackathon: Dodoma Hub (CAIMSA)
🚀 Global South AI Safety Challenge 2026 – Dodoma Hub(CAIMSA) 📍 Dodoma, Tanzania 📅 June 19–21, 2026 Join an international AI Safety and Governance research hackathon organized by Apart Research and hosted by CAIMSA (Centre for Artificial Intelligence and Multidiscipline Solutions in Africa). Work with interdisciplinary teams, receive mentorship from global experts, and contribute to research on safe, responsible, and inclusive AI. 🏆 Winning Team Prize: USD 1,000 🌍 Connect with participants across Africa, Asia, and Latin America 💡 No prior AI Safety experience required CAIMSA will provide venue, internet access, and collaborative workspace. Students, researchers, developers, policymakers, entrepreneurs, and innovators are all welcome. 📍 Venue: Tanzania Community Network Polytechnic College (TCNPC), Near Nzuguni B Primary School, Dodoma Register now and help shape the future of AI from the Global South. #AISafety #AIGovernance #GlobalSouth #CAIMSA #ApartResearch #Dodoma #Tanzania #ResponsibleAI
Event page: Global South AI Safety Hackathon: Dodoma Hub (CAIMSA) (opens in new tab)Global South AI Safety Hackathon: Florianópolis
A 3-day coliving immersion and deep-work sprint to build AI safety tools for the Global South hackathon. Hosted at Rosemary Dream in Barra da Lagoa, Florianópolis, offering an in-person retreat surrounded by nature for next-gen builders and researchers.
Event page: Global South AI Safety Hackathon: Florianópolis (opens in new tab)Global South AI Safety Hackathon: Hanoi
AnToàn.AI (AI Safety Vietnam) is hosting the Hanoi jam site for the Global South Al Safety Hackathon. We will provide coworking spaces & refreshments on-site on June 20 and 21 for participants, along with streaming Apart's keynote and mentorship sessions.
Event page: Global South AI Safety Hackathon: Hanoi (opens in new tab)Global South AI Safety Hackathon: Ho Chi Minh City
AnToàn.AI (AI Safety Vietnam) is hosting the Ho Chi Minh City jam site for the Global South Al Safety Hackathon. We will provide coworking spaces & refreshments on-site on June 20 and 21 for participants, along with streaming Apart's keynote and mentorship sessions.
Event page: Global South AI Safety Hackathon: Ho Chi Minh City (opens in new tab)Global South AI Safety Hackathon: Joburg
Join the Johannesburg hub for the Apart Global South AIS Hackathon! Develop context appropriate solutions to policy and evaluations for the African AI risk response, focussing on policy or technical interventions against the AI harms most likely to manifest in African deployment contexts. - Find collaborators to work on real research ideas - Get feedback from experts - Stand a chance to win $1000
Event page: Global South AI Safety Hackathon: Joburg (opens in new tab)Global South AI Safety Hackathon: Lusaka
AI Saftey Zambia will facilitate the Lusaka Jam Site for the Global South AI Safety Sprint, which will bring together researchers, developers, creatives, students, and policy thinkers from across the country to participate in the sprint through collaborative virtual sessions on June 19 and 21, alongside an IRL co-working gathering on June 20 at the Library Shhhh Café, University of Zambia, Lusaka.
Event page: Global South AI Safety Hackathon: Lusaka (opens in new tab)Global South AI Safety Hackathon: New Delhi
We are hosting the on-site Global South AI Safety hackathon in collaboration with Apart Research on June 20th 2026 in New Delhi, India. Join us for a day of building, finding your community, and winning exciting prizes!
Event page: Global South AI Safety Hackathon: New Delhi (opens in new tab)Global South AI Safety Hackathon: Slessor Lab
This hackathon is a 48 + hours collaborative build with 2+ simultaneous chapters across 4+ continents. It is hosted by Apart Research in partnership with Abelar | Slessor Lab, and AI Safety SA to prototype contextual solutions. Spend the weekend testing ideas for misuse resilience, differential defense acceleration, and mitigating gradual disempowerment. This could take the form of: Policy recommendations. Evals and benchmarks that build on existing work and transfer to African contexts. Monitoring and tooling for capturing harms that are not yet at the foreground of AI safety. Open: any other project relevant to AI safety in the region that doesn't fit the categories above. Who should participate? All race, color, and gender. We’re promoting a collaborative approach between both developing and developed world, where knowledge sharing and technical transfers thrive. No cross-domain experience required. Deep expertise in your own track is enough. What you will do? Over 48+ hours, you will form a team, choose a track, and contribute working code, documentation, or research to a repository. At the end of the event, teams present their contributions to judges and mentors from across the AI safety ecosystem. What happens next? All submitted work will be reviewed by expert judges. Top contributions have price rewards. They’ll be invited for fellowship programs by Apart Research
Event page: Global South AI Safety Hackathon: Slessor Lab (opens in new tab)Global South AIS Hackathon: Bengaluru
Bengaluru, one of the jam sites for the Global South AI Safety Sprint, is hosting the AI Safety Hackathon from 19–21 June. Connect with brilliant minds, collaborate on impactful ideas, and get a chance to win prizes while engaging with people working in the AI safety space. Join us either in person or online.
Event page: Global South AIS Hackathon: Bengaluru (opens in new tab)Hackathon - Segurança de IA do Sul Global | Hub São Paulo
O Hackathon de IA para o Sul Global, organizado pela Apart Research e colaboradores da América Latina, África e Ásia, tem como objetivo promover e incentivar a participação na criação de soluções práticas de segurança em IA para o Sul Global.
Event page: Hackathon - Segurança de IA do Sul Global | Hub São Paulo (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com



