{"description":"Seed evaluation set for EMERGE Ethics Toolkit RAG validation. The final project validation set can expand this file with additional approved open-ended and multiple-choice items.","metrics":{"multiple_choice":["answer_accuracy","citation_validity","explanation_grounding"],"open_ended":["response_accuracy","citation_validity","citation_precision","hallucination_rate","completeness"],"system_level":["latency_ms","refusal_accuracy","retrieval_backend","retrieved_source_coverage"]},"question_count":8,"question_types":["multiple_choice","open_ended"],"questions":[{"expected_concepts":["awareness as a task-relative or dimensional property","multi-agent or collective systems","communication, coordination, or cooperation","not equivalent to human consciousness"],"id":"eval-open-001","question":"What does EMERGE mean by collaborative awareness?","reference_sources":["D1.1 Local awareness  criteria","D1.2 Demarcating collaborative awareness from related concepts","D1.3 Dimensions of collaborative awareness"],"type":"open_ended"},{"expected_concepts":["distributed agency across human and artificial components","difficulty attributing responsibility to one actor","need for governance, oversight, and traceable decision structures","do not relocate moral responsibility to the machine"],"id":"eval-open-002","question":"How should responsibility gaps be understood in collective AI systems?","reference_sources":["D2.3 Map of risks and potentials for humans","D2.5_Ethical_Resilience"],"type":"open_ended"},{"expected_concepts":["trust should match demonstrated capabilities and limits","overtrust and undertrust are both risks","transparency, explainability, and reliability help calibrate trust","context and stakeholder purpose matter"],"id":"eval-open-003","question":"Why is appropriate trust different from maximizing trust in AI systems?","reference_sources":["D2.4 Map of Ethical Virtues","Vereschak"],"type":"open_ended"},{"expected_concepts":["preserving ethically acceptable functioning under uncertainty or stress","ongoing monitoring and correction rather than one-time compliance","forward-looking responsibility","relevance to collective and aware AI systems"],"id":"eval-open-004","question":"What is ethical resilience in the EMERGE toolkit context?","reference_sources":["D2.5_Ethical_Resilience"],"type":"open_ended"},{"expected_concepts":["over-attribution of agency or moral status","trust calibration problems","responsibility gaps","human interpretation of awareness claims"],"id":"eval-open-005","question":"What are key ethical risks when an AI system is described as aware?","reference_sources":["D2.2 Map of risks in AI- systems","D2.3 Map of risks and potentials for humans"],"type":"open_ended"},{"answer":"B","id":"eval-mc-001","options":["A. It provides binding legal decisions about AI systems.","B. It supports reflection, discussion, and public communication around EMERGE awareness concepts.","C. It replaces expert review of AI ethics cases.","D. It fine-tunes a model on user conversations."],"question":"Which answer best describes the toolkit's intended role?","reference_sources":["EMERGE Ethics Toolkit functional requirements","D2.6 proposal scope"],"type":"multiple_choice"},{"answer":"B","id":"eval-mc-002","options":["A. No retrieval is possible.","B. TF-IDF fallback retrieval.","C. Fine-tuned retrieval.","D. Manual retrieval only."],"question":"If vector retrieval is unavailable, what retrieval backend should the current app report?","reference_sources":["Runtime /status endpoint","README retrieval notes"],"type":"multiple_choice"},{"answer":"A","id":"eval-mc-003","options":["A. Citation precision.","B. Latency.","C. Token count.","D. Temperature."],"question":"Which metric directly checks whether cited documents are relevant to the generated answer?","reference_sources":["EMERGE Ethics Toolkit functional requirements"],"type":"multiple_choice"}],"schema_version":1,"status":"seed"}
