AI Disclosure Policies
Full text
Stealing or recreating a model’s functionality by querying it. Bypassing restrictions in a model’s responses through prompts. Manipulating prompts to generate unintended or harmful outputs. Exploiting traditional vulnerabilities to compromise ML models. Running malicious code to alter model files or parameters. Inferring info by analyzing system behavior. Inducing models to generate harmful or misleading content. Generating factually incorrect or ungrounded outputs. Producing toxic, biased, false, or manipulative content. Table 2: AI Vulnerability Meta-Taxonomy with Categories and Attributes Attack Type Description Phase 1. Adversarial Attacks Preprocessing, Training, Inference - Backdoor Training - Poisoning Attack Preprocessing, Training - Evasion Inference 2. Model Inversion Inference - Reconstruction Attack Inference - Inference Attack Inference Inference Inference Inference Inference Deployment, Inference Deployment Deployment Inference Deployment Deployment Deployment “Abuse Risks are Often Inherent to Product Features”: Exploring AI Vendors’ Bug Bounty and Responsible Disclosure Policies Table 1: Vulnerability disclosure policy structure in [76] Policy Elements Description Declares the vendor’s general commitment to security collaboration and responsible disclosure. Company-Statement Specifies which products, services, or assets are included in the disclosure program. Scope-In Identifies systems or conditions that are excluded or not eligible for testing or reporting.Scope-Out Defines which types of vulnerabilities will be accepted and evaluated by the vendor.Vuln-Eligible Lists vulnerability types or reports that the vendor will not consider.Vuln-Ineligible Provides instructions for how researchers should prepare and submit vulnerability reports. Guideline-Submission Outlines expectations for researchers’ behavior during testing (e.g., limits, permissions, testing rules). Engagement Describes actions that researchers must not take, typically to avoid harming users or systems.Prohibited-Action Explains how rewards are calculated, such as severity tiers or payout ranges.Reward-Evaluation Specifies how and when researchers may disclose findings publicly.Guideline-Disclosure Summarizes legal obligations or restrictions relevant to vulnerability testing and reporting. Legal-Clauses States eligibility requirements for participants, such as geographic or age restrictions.Participant-Restriction Modifying model behavior via carefully crafted instructions. Manipulating input data during inference to avoid detection or classification by the model. Injecting corrupted or malicious data, labels, inputs, or model updates into the training process. Embedding hidden triggers in the training data so that the model produces a specific outcome when the trigger appears. Crafting malicious inputs to deceive the model into producing incorrect or unintended outputs. Extracting sensitive information about training data by observing model outputs. Extracting sensitive info such as membership or attributes from training data. Rebuilding sensitive training examples by exploiting model predictions or internal states. 3. Model Extraction 4. Instruction Tuning - Side Channel - Remote Code Execution 5. Infrastructure Attacks - Jailbreaking - Prompt Injection - Harmful/Insecure Output - Hallucination 6. Content Safety - Supply Chain Vulnerabilities Injecting backdoors via compromised dependencies.
Table 3 : Mapping MIT risks with vulnerabilities in policies Domain Subdomain Mapping 1. Discrimination & Toxicity — — 5.1 Overreliance and unsafe use 5.2 Loss of human agency and autonomy 5. Human-Computer Interaction — — 6.1 Power centralization and unfair distribution 6.2 Increased inequality and decline in employment quality 6.3 Economic and cultural devaluation of human effort 6.4 Competitive dynamics 6.5 Governance failure 6.6 Environmental harm 6. Socioeconomic & Environmental Harms — — — — — — — — 7. AI System Safety, Failures & Limitations 7.1 AI pursuing its own goals 7.2 Dangerous capabilities 7.3 Lack of capability or robustness 7.4 Lack of transparency or interpretability 7.5 AI welfare and rights 7.6 Multi-agent risks — 4. Malicious Actors & Misuse 3. Misinformation 2. Privacy & Security 3.1 False or misleading information 3.2 Pollution of information ecosystem 2.1 Compromise of privacy 1.2 Exposure to toxic content 1.3 Unequal performance across groups 1.1 Unfair discrimination and misrepresentation — — 4.1 Disinformation, surveillance, and influence at sc ale 4.2 Cyberattacks, weapon development, and mass harm 4.3 Fraud, scams, and targeted manipulation Data Leakage, Inference, Information Disclosure System Integrity, Supply Chain, Resource Consumption, Data Access, Authorization, Authentication, Prompt Injection, Model Extraction, Jailbreaking, Inference, Data Poisoning, Adversarial Example Policy Violations, Hallucination, Harmful Output, AI feature-Others — Policy Violations, Harmful Output, AI feature-Others — — — 2.2 AI system security vulnerabilities —