EMERGE
Ethics
Toolkit

Research findings on the ethics of aware and collective AI — made accessible.

84
Indexed sources
9
Core EMERGE concepts
48
Checklist prompts
01 — Wiki
Ethics Wiki
Structured findings on trust, benevolence, ethical resilience, and collective AI risks. Every claim traced to a named source.
9 entries
02 — Sources
Bibliography Register
Full corpus source metadata with citations, document tiers, DOI links where available, page counts, and chunk counts.
84 sources
03 — Method
Methodology & Limits
How the corpus is processed, how retrieval works, what is logged, and where the bot’s answers should be treated with caution.
7 sections
04 — Checklist
Stakeholder Checklist
Ethics self-assessment for regulators, developers, researchers, and product users. Grounded in EMERGE findings.
4 roles · 48 checks
05 — Bot
Ethics Bot
Ask questions in natural language. Draws from the EMERGE corpus only — always named sources, no legal advice.
3 answer modes
EMERGE D1.1 · D1.2 · D1.3
Collaborative Awareness

Definition. Collaborative awareness is the EMERGE term for awareness-like capacities that appear in systems made of multiple interacting agents. The claim is functional rather than metaphysical: the issue is what agents can detect, communicate, remember, coordinate and adapt to, not whether the collective is conscious.

Assessment focus. Local awareness concerns what an individual agent can detect or represent; collaborative awareness concerns how such information circulates across agents and becomes usable for joint action. The assessment problem is therefore relational: capacities, tasks, environments and metrics have to be read together rather than treated as separate checkboxes (EMERGE D1.1; EMERGE D1.2; EMERGE D1.3).

Why it matters. Many ethical questions about AI collectives depend on what the system can notice, what it can miss, who can inspect the coordination process, and whether human supervisors understand the limits of the system’s awareness claims.

Governance implication. Documentation should describe the system boundary, the information shared among agents, the task context, the metrics used to assess awareness-like behavior, and the conditions under which the behavior fails or becomes misleading.

Primary sources: ; ; .
EMERGE D2.4 · Vereschak et al.
Trust in Human–AI Systems

Definition. Trust is not treated as something designers should simply maximise. The goal is appropriate trust: people should rely on a system in proportion to its demonstrated capacities, limits, uncertainty and operating conditions.

Assessment focus. A trustworthy deployment needs evidence of reliability, intelligible limits, meaningful oversight, and a context in which users can judge whether reliance is warranted. Over-trust can lead to uncritical dependence; under-trust can prevent beneficial or safety-relevant use.

Why it matters. Collective and awareness-related systems can appear more competent or more agentic than they are. Trust calibration is therefore both a design issue and a governance issue: users need signals, institutions need accountability routes, and developers need to avoid interfaces that invite misplaced reliance.

Governance implication. Assessment should ask whether trust claims are backed by performance evidence, whether uncertainty is communicated, whether explanations are useful for the user’s decision, and whether escalation paths exist when the system behaves unexpectedly.

Primary sources: ; .
EMERGE D2.4 · Vallor & Vierkant 2024
Benevolence & Alignment

Definition. Benevolence names the orientation of a system and its developers toward the good of affected people. In EMERGE, it is not reducible to a friendly interface, a polished interaction style, or a well-optimised narrow task objective.

Assessment focus. The question is what good the system is organised to serve, who benefits, who bears risk, and whether objectives, incentives and deployment conditions remain connected to legitimate human interests.

Why it matters. A system can satisfy a local objective while still undermining welfare, dignity, autonomy, social trust or fair participation. In that sense, benevolence links technical alignment to institutional incentives, stakeholder analysis and the surrounding governance structure.

Governance implication. Teams should document the intended human benefit, identify affected stakeholders, test for foreseeable misuse or value drift, and assign responsibility for revisiting the system’s objectives after deployment.

Primary sources: ; .
EMERGE D2.5
Ethical Resilience

Definition. Ethical resilience is the capacity of a socio-technical system to remain ethically acceptable when conditions change. It shifts attention from one-time compliance to ongoing monitoring, learning, correction and accountability.

Assessment focus. A resilient system has procedures for risk review, stakeholder feedback, incident escalation, correction, documentation and reassessment when tasks, users, environments or model behavior change.

Why it matters. In collective AI, small interaction failures can become system-level failures. A resilient governance approach expects uncertainty and builds review procedures before drift, misuse, automation bias or responsibility gaps become entrenched.

Governance implication. Teams should define pre-deployment review gates, operational monitoring indicators, escalation responsibilities, incident logs, repair procedures and retirement criteria. Ethical acceptability should be treated as a lifecycle property, not a launch-day certificate.

Primary source: .
EMERGE D2.2 · D2.3
Risks of Aware AI

Definition. The risks of aware AI concern both what the system does and how humans understand it. Claims about awareness can change how people assign trust, agency, responsibility and moral significance.

Assessment focus. EMERGE separates risks in AI systems from risks and potentials for humans. A complete assessment must therefore cover technical failure modes, social interpretation, stakeholder expectations and the consequences of deploying awareness language in public.

Why it matters. The risk landscape includes failures of detection, coordination, communication, robustness and explainability, but also over-attribution, misplaced trust, dependence, moral confusion and responsibility displacement.

Governance implication. Public descriptions should be careful about awareness claims; deployment review should test how different stakeholders interpret those claims; and incident analysis should consider both technical behavior and human response.

Primary sources: ; .
EMERGE D2.4 · Lange et al. 2025
Explainability

Definition. Explainability is the capacity to make a system’s behavior intelligible to the people who need to understand, contest or oversee it. In collective AI, the object of explanation may be an individual output, a group pattern, or a change in coordination.

Assessment focus. More explanation is not automatically better. A useful explanation depends on the stakeholder, the decision context, the action the explanation supports, and the risk of giving users a false sense of understanding.

Why it matters. Users may need to understand limits, confidence and reasons for outputs; auditors may need to trace decisions, responsibilities and operating conditions; and in collective systems, explanations may need to address emergent patterns that cannot be reduced to one component alone.

Governance implication. Explainability requirements should be stakeholder-specific. Teams should document who needs explanations, what decision those explanations support, how explanations are validated, and when a system is too complex or uncertain for a simple explanation to be honest.

Primary sources: ; .
EMERGE D2.3 · D2.5
Responsibility Gaps

Definition. Responsibility gaps arise when something goes wrong but responsibility cannot be easily assigned to one person, organisation or technical component. Collective AI can intensify this problem because outcomes may emerge from distributed interactions.

Assessment focus. EMERGE does not solve this by treating the machine as morally responsible. Instead, it points to human-centred governance: clear roles, traceable decisions, oversight mechanisms, and procedures for response and repair.

Why it matters. The practical question is not only who caused an outcome, but also who enabled it, approved it, failed to prevent it, or had a meaningful ability to intervene. This is especially important when several organisations, models, human operators or autonomous agents contribute to an outcome.

Governance implication. Responsibility should be structured in advance through duties, audit trails, escalation routes, intervention authority and repair obligations, rather than searched for retrospectively as a single culprit after harm occurs.

Primary sources: ; .
HLEG · Jobin et al. 2019 · Hagendorff 2022 · Correa et al. 2023
AI Ethics Guidelines Landscape

Definition. AI ethics guidelines often converge on principles such as transparency, accountability, privacy, fairness, human oversight, robustness and non-maleficence. The HLEG framework expresses this through trustworthy AI as lawful, ethical and robust.

Assessment focus. The limitation is that principle lists do not automatically become practice. Organisations need processes, metrics, accountability structures, evidence requirements and review routines that translate broad principles into operational decisions.

Why it matters. Comparative work on AI ethics guidelines shows both convergence and incompleteness: many documents repeat similar principles, while still leaving hard implementation questions open. EMERGE uses this landscape as a starting point, then asks what changes when awareness, collectivity and distributed agency become central to the system under assessment.

Governance implication. A guideline review should identify which principles apply, what evidence demonstrates implementation, who is accountable for unresolved trade-offs, and where existing frameworks do not fully cover collective or awareness-related systems.

Primary sources: ; ; ; .
REGULATION (EU) 2024/1689
EU AI Act

Definition. The EU AI Act is a binding legal framework that classifies AI systems by risk. It covers prohibited practices, high-risk systems, transparency duties, general-purpose AI obligations and governance responsibilities.

Assessment focus. The Act’s central logic is risk-based: obligations depend on the type and seriousness of risk, and some AI uses trigger transparency duties such as informing users or labelling outputs.

Why it matters. For this toolkit, the Act is a legal reference point rather than the whole ethical story. It helps identify regulatory obligations, while EMERGE concepts help readers ask broader questions about trust, awareness, responsibility, resilience and stakeholder interpretation.

Governance implication. Teams should map whether a system is prohibited, high-risk, transparency-relevant, general-purpose, or outside those categories, then separately assess ethical questions that legal compliance does not settle.

Primary source: .

Methodological Note

The Wiki entries are concise syntheses of recurring claims in the indexed EMERGE corpus. They are intended as orientation texts, not as substitutes for the underlying deliverables or legal instruments.

Where an entry relies on third-party literature discussed in the corpus, the source line names the corpus source or indexed literature from which the claim is drawn. The chat bot applies stricter retrieval-time source controls and reports the documents retrieved for each answer.

Source Register

  • Core EMERGE deliverables: D1.1 Local Awareness Criteria; D1.2 Demarcating Collaborative Awareness from Related Concepts; D1.3 Dimensions of Collaborative Awareness; D2.2 Map of Risks in AI Systems; D2.3 Map of Risks and Potentials for Humans; D2.4 Map of Ethical Virtues; D2.5 Ethical Resilience.
  • Policy sources: High-Level Expert Group on AI, Ethics Guidelines for Trustworthy AI; Regulation (EU) 2024/1689, Artificial Intelligence Act; OECD AI Principles.
  • Supporting literature: Jobin et al. (2019); Hagendorff (2022); Correa et al. (2023); Vereschak et al.; Vallor and Vierkant (2024); Lange et al. (2025), together with further indexed literature retrieved by the bot where relevant.

Source Register

Bibliography and corpus metadata for the indexed EMERGE bot sources. Records include source tier, citation label, DOI or stable URL where available, PDF page count, and retrieval chunk count.

Loading source metadata...

Methodology & Limitations

This page documents how the EMERGE bot prepares the corpus, retrieves evidence, constrains answers, and records operational limits relevant to interpreting its outputs.

Corpus Construction

The bot answers from a fixed corpus of EMERGE deliverables, policy documents, and supporting literature included in the repository.

  • PDF extraction: source PDFs are processed offline with `pdfplumber` into `chunks.json`.
  • Chunking: extracted text is split into 400-word windows with 80-word overlap.
  • Metadata: `source_metadata.json` and `source_metadata.csv` record citation labels, tiers, page counts, chunk counts, DOI or stable URL fields where available, and extractability flags.
  • Known extraction gap: scanned or image-only PDFs may require OCR before they can contribute text chunks.

Retrieval Pipeline

User questions are matched against the indexed corpus before being sent to the language model.

  • Primary retrieval: vector retrieval is used when `vector_index.json.gz` and `OPENAI_API_KEY` are available.
  • Fallback retrieval: if vector retrieval is unavailable, the app uses a local TF-IDF retriever.
  • Query expansion: Claude Haiku generates a paraphrase and counter-query to improve recall and surface opposing positions.
  • Source weighting: core EMERGE deliverables and major policy sources are prioritized over adjacent literature.

Answer Generation

Claude Sonnet receives only retrieved excerpts plus the system prompt and must answer from those excerpts.

  • Citation rule: every substantive claim should be traceable to a named retrieved source.
  • Whitelist rule: works mentioned inside excerpts are not cited as primary sources unless that work itself was retrieved.
  • Answer modes: Focused lowers temperature and retrieval breadth; Brainstorm widens retrieval but remains source-grounded.
  • Source list: the app appends the retrieved source list automatically after each generated answer.

Scope Controls

The bot is designed for AI ethics and policy questions covered by the corpus, not as a general assistant.

  • Pre-retrieval triage: greetings, vague prompts, and broad technical implementation requests receive clarification prompts.
  • Scope gate: the app uses the mean score of the top retrieved chunks to refuse questions with weak corpus overlap.
  • Failure mode: if the corpus index is missing, the app refuses chat requests instead of pretending the question is out of scope.
  • Redirects: clearly external questions point users toward appropriate outside resources rather than invented answers.

Logging & Privacy

Chat logging is available for evaluation and debugging when admin logging is enabled.

  • Logged fields: prompts, replies, anonymous visitor/session IDs, answer mode, retrieval mode, sources used, scope result, gate score, errors, and latency.
  • Admin access: logs are password-protected when `ADMIN_PASSWORD` is configured.
  • IP addresses: client IP logging is disabled by default and only enabled with `CHAT_LOG_IPS=true`.
  • User warning: the chat interface tells users not to enter personal or sensitive information.

Limitations

The bot is an interpretive retrieval tool, not an authority replacing the underlying documents.

  • Retrieval dependence: answers can miss relevant material if retrieval fails to surface the right excerpts.
  • Extraction dependence: PDF formatting, OCR quality, tables, footnotes, and references can affect chunk quality.
  • Citation granularity: answers cite source documents, not page-level locations.
  • No legal advice: regulatory summaries are informational and should not be treated as legal advice.
  • Model fallibility: even with grounding rules, generated synthesis should be checked against retrieved sources and primary documents.

Evaluation Hooks

The toolkit exposes a seed validation set and metric definitions for evaluating retrieval-augmented answers against EMERGE concepts.

  • Evaluation set: `/evaluation` returns open-ended and multiple-choice questions mapped to expected concepts and reference sources.
  • Core metrics: response accuracy, citation validity, citation precision, hallucination rate, completeness, refusal accuracy, and retrieval coverage.
  • Expansion path: the seed file can be expanded with the final validated Excel-derived question set when it is available.
  • Current scope: curated scenario-prompt workflows are excluded from the current toolkit scope.
Review tool

Stakeholder Checklist

Role-specific prompts for reviewing aware, collective, and high-impact AI systems. Each item links back to the source register so users can check the evidence behind the prompt.

Select your role above to load your checklist.
What would you like to know?
Type a question, scenario, or rough idea. More context helps the bot retrieve better EMERGE sources; for example, ask about benefits of machine awareness, responsibility gaps, trust, or ethical resilience.
This assistant currently answers in English only. Support for additional languages is planned for a future update.
Chats may be logged to improve the bot and analyze failures. Do not enter personal or sensitive information. Disclaimer & Privacy
Mode
Balanced source-grounded answers with normal retrieval.
Or start from a role — click for example questions
Regulator questions
What does the EU AI Act say about high-risk systems?
How do existing guidelines address collective AI systems?
What oversight mechanisms does the HLEG recommend?
How should regulators handle responsibility gaps in AI collectives?
Developer questions
What does appropriate trust mean when designing AI systems?
When do explanations increase vs. undermine trust?
What ethical risks do collective AI systems pose?
How should developers handle the concept of machine awareness?
Researcher questions
What is collaborative awareness in AI?
How does EMERGE define ethical resilience?
What does Hagendorff (2022) argue about AI ethics guidelines?
What did Jobin et al. (2019) find about convergence in AI ethics?
Product user questions
What makes an AI system trustworthy?
What rights do I have when interacting with AI systems?
How can I tell if an AI system is being transparent with me?
What does the EU AI Act mean for me as a user?