Evolutionary Computation for Trustworthy AI: From Attack-Defense Search to Self-Evolving Systems
This review examines how evolutionary computation supports trustworthy AI across three research lines: evolutionary attacks to find failures, evolutionary defenses to search protective measures, and trustworthy self-evolution to govern which changes persist long-term, highlighting the shift from static robustness to continuous governance of memory, skills, and safeguards.
Introduction
The review interprets the arXiv preprint "Evolutionary Computation for Trustworthy AI: From Attacks and Defenses to Self-Evolving Era" (arXiv:2610.02996, submitted Oct 2, 2026, 20 pages). Authors: Junhao Dong, Chenkai Wang, Xuanhui Lin, et al. (first three co-first). Affiliations include Nanyang Technological University, A*STAR, Google, Sichuan University, Lingnan University.
Trustworthy AI evaluation objects have expanded from classifiers to foundation models, agents, and long-running systems. Trustworthiness now covers tool use, external knowledge, memory, and continuous updates. These objects are often discrete, structured, non-differentiable, and expensive to evaluate; robustness, safety, task utility, and cost conflict. Single-gradient or single-score methods cannot cover these challenges.
The survey connects three research threads through evolutionary computation: evolutionary attacks search for failures, evolutionary defenses search for protections, and trustworthy self-evolution governs which changes can be inherited long-term. A key shift: systems must decide which memories, skills, workflows, and safeguards enter the next generation. An insufficiently verified "improvement" may alter future behavior and future update generation.
Survey Scope and Positioning
Scope
Focuses on the intersection of evolutionary computation with robustness, safety, security, and privacy. Fairness appears in multi-objective design but the survey is not a comprehensive governance review. Three roles of evolution: attack side searches perturbations, triggers, prompts, knowledge, or policies; defense side optimizes data, structure, training, and deployment protections; self-evolution side changes persistent state available to future versions. The survey includes both standard evolutionary algorithms and evolutionary-inspired systems that repeatedly generate, evaluate, retain, and reuse memories and skills.
Comparison with Existing Surveys
Existing surveys organize trustworthy AI by threat models/attack surfaces/defenses; evolutionary learning by representation/mutation/selection/multi-objective/knowledge transfer; self-evolution by long-term changes to experience/memory/workflows. This survey links them, highlighting the common question: how candidates are generated and evaluated, and whether selected results become part of future systems. Table 1 and Figure 3 (2018-2026 representative works) are the authors' literature organization, not a unified benchmark ranking.
Preliminaries
Evolutionary Search Process
Standard flow: initialize population, select parents by fitness, produce offspring via mutation/recombination, decide next-generation retention until budget/stop condition. Fitness can be scalar or vector (safety, utility, cost). Memetic computing combines population search with local improvement (gradients/task knowledge) without losing diversity. Artificial immune systems borrow recognition, adaptive defense, immune memory for detection/response. Thus "no analytical gradients" does not forbid gradient use in hybrids.
Attacks and Defenses via Evolutionary Search
Traditional adversarial learning: attack maximizes loss within allowed perturbation set; defense minimizes worst-case expected loss. Evolutionary computation expands search to candidate populations, not single perturbations or model parameters. Multi-objective populations preserve trade-offs (effectiveness, stealth, cost for attacks; robustness, utility, efficiency for defenses). Quality-diversity archives keep behaviorally complementary candidates (e.g., triggering different failure types). When attack and defense co-update, one side's improvement changes the other's fitness landscape, forming co-evolution.
State Updates in Self-Evolving Systems
System state includes model components, memory, skills, workflows, safeguards, archives. Candidate updates undergo admission judgment: if passed, state changes; otherwise original state retained. Accepted state is inherited next round, influencing subsequent candidate generation, evaluation, selection. This separates "generating candidates" from "approving long-term writes". Trustworthiness can enter at three points: constrain allowed update space; guide selection/admission; shape continuous attack-defense feedback. Verifying one output safe does not prove writing related experience to long-term memory is safe.
Evaluation and Benchmarking
Trustworthiness Evaluation
Robustness
Attack success rate must be interpreted under explicit threat models. Images report perturbation magnitude and query cost; sparse/hard-label attacks involve modified element count and boundary search cost. Text, speech, graphs require semantic, perceptual, structural validity constraints. High success rate with altered task meaning may not indicate valid robustness test.
Safety and Privacy
Safety evaluation looks beyond harmful response ratio to exposed failure types. Quality-diversity red-teaming reports behavioral coverage and archive quality. Model inversion involves reconstruction similarity, identity-related success rate, transferability. Privacy evaluation must distinguish membership inference, training data memorization, etc.; single attack metric cannot replace all privacy conclusions.
Task Performance Retention
Clean accuracy, task success rate, normal request refusal rate should be reported alongside safety metrics. Over-refusal reduces utility; reducing refusal may relax safety boundaries. Trustworthy optimization should present trade-offs, not just higher refusal or lower attack success.
Evolutionary Search Evaluation
Optimization Performance and Diversity
Single-objective: best/final fitness. Multi-objective: hypervolume for trade-off sets. Quality-diversity: highest fitness, archive coverage, QD score. Coverage only meaningful if behavior categories are well-defined.
Budget and Convergence
"One evaluation" may be a model query, judge call, architecture training, or environment execution — costs not equivalent. Fixed-budget comparison: results under same resources. Fixed-target comparison: cost to reach specified level. Paper emphasizes independent repeated runs, distribution statistics, confidence intervals to avoid mistaking lucky runs for stable advantage.
Benchmark Resources
Resources split into trustworthiness benchmarks (RobustBench, NAS-RobBench-201, JailbreakBench, HarmBench, StrongREJECT, MM-SafetyBench, ETHICS, BBQ, Agent-SafetyBench, R-Judge, OS-Harm) and evolutionary search benchmarks (COCO/BBOB, IOH, multi-objective test suites, QD-Suite, QDax, multi-task optimization resources). Full comparison requires specifying target system, threat model, evaluator, task metrics, cost unit, total budget, stop condition, independent runs. These resources provide measurement tools but do not yet form a standardized unified leaderboard.
Evolutionary Attacks
Attacks on Predictive Models
Adversarial Input Generation
One Pixel Attack, GenAttack treat perturbations as population candidates, using model feedback to search failures. Extended to high-dimensional, sparse, label-only black-box conditions; cover text, speech, graph data. Common value: handling non-differentiable or limited-feedback search while satisfying semantic/structural/functional constraints.
Evolutionary Backdoor Attacks
Backdoor research moves search to persistent manipulation during training. LADDER optimizes trigger patterns balancing effectiveness and stealth; others explore frequency-domain patterns and input-dependent triggers. Trustworthiness analysis must distinguish inference-time perturbations from training-time implants because protection locations and persistence differ.
Evolutionary Model Inversion
Inversion aims not just to change predictions but to use output feedback to approach target-identity information. Survey discusses genetic search combined with generative priors; evaluation includes reconstruction quality, similarity, relevant success rates. This shows evolutionary search's link to privacy risk, not merely robustness metric extension.
Attacks on Generative Models
Evolutionary Jailbreaking
Open Sesame, AutoDAN treat prompts as structured candidates, using selection, hierarchical recombination, LLM-assisted mutation. Later work adds semantic quality, stealth, cross-model transfer objectives. Focus is on search mechanisms and evaluation design, not providing reusable jailbreak texts.
Retrieval-Augmented Attacks and Model Extraction
RAG introduces external knowledge channel independent of parameters. GARAG, DIGA, NeuroGenPoisoning search documents/passages to bias retrieval and generation; Stealix searches informative queries to improve extraction efficiency. These manipulate external evidence or extract model behavior, not all reducible to prompt jailbreaking.
Attacks on Agent Systems
Injection and Interface Attacks
Agents read from web pages, tables, UIs, tool responses, expanding input boundaries. StruPhantom, Genesis, EVA inject candidates into these interfaces, evaluating impact on subsequent agent behavior. Risk includes induced operations and process deviations, not just bad answers.
Trajectory and Environment-Aware Attacks
T-MAP uses tool execution trajectories to evaluate candidates and guide search; OpenART incorporates environment state into open-ended red-teaming. Evaluation object expands from single input to chain of actions, observations, environment changes. Long-term memory, multi-agent interaction, embodied execution remain relatively underexplored.
Discussion
Attack effectiveness depends on evolutionary algorithm, representation, initialization, feedback quality, fitness, budget. Existing work often changes multiple factors simultaneously, making it hard to isolate the evolutionary mechanism's contribution. Authors suggest more controlled black-box search comparisons and efficient evaluators capturing retrieval, tool, and long-trajectory propagation effects.
Evolutionary Defenses
Training Data
Evolutionary Adversarial Training
Defense feeds discovered hard samples back to training. EMO-GAT searches sample selection and proportions; ER-APT combines gradient generation with evolutionary operations to construct diverse samples for vision-language model adversarial prompt learning; IGAff searches transformed hard images. Evolution mainly changes the data distribution the model learns from.
Safety and Alignment Data Generation
EVOREFUSE targets unnecessary refusals from pseudo-malicious instructions; Rainbow Teaming uses quality-diversity search to generate diverse failure prompts for subsequent training; Q-DIG extends similar patterns to vision-language-action models. Distinction: discovering failures vs. using failures for training — the latter requires verification of improvement on unseen scenarios.
Model Design
Robust Architecture Search
Evolutionary architecture search incorporates clean performance, adversarial robustness, model size directly into structure selection. MORAS, dual-fidelity search, TAM-NAS, NERO-Net represent different candidate spaces and cost strategies. Module switching searches combinations of existing models to suppress backdoors while retaining utility. Model design can also explicitly optimize ensemble diversity and accuracy-fairness trade-offs. Core is not "certain structures are inherently safe" but incorporating trustworthiness into structural evaluation and continuously verifying performance across attacks and data distributions.
Training Optimization
Objective and Training Strategy Search
TaylorGLO evolves loss functions; MO-PBT searches multi-objective training configurations; TRG-ASO adjusts adversarial strategies across training stages. EvoPref maintains low-rank adapter populations exploring trade-offs among helpfulness, harmlessness, honesty. Search objects can be training rules or alignment parameters, not just input samples.
Safety and Robust Policy Search
SNES, SMC-ES combine evolution strategies with model checking; SERL studies fault-tolerant control; LyEvO integrates constrained optimization, statistical verification, stability analysis. Explicit safety requirements can enter fitness or screening, but verification conclusions always depend on specified behavioral properties and environment assumptions.
Deployment Protection
Evolutionary Inference-Time Defense
RAILS uses immune-inspired process to adjust class-balanced sample populations near current input; safe face recognition works search input-specific denoising strategies. These place adaptation inside inference loop, requiring measurement of extra time and query overhead, not just final robustness.
Post-Training Backdoor Mitigation
Trigger region detection and lightweight repair use evolutionary search to locate suspicious regions then repair; module switching optimizes post-training configurations. Effect must jointly measure backdoor suppression and normal performance; detecting candidate regions should not be equated with proven removal of all backdoors.
Discussion
Evolutionary computation typically acts as outer optimizer, not replacing underlying learning and safety mechanisms. Key questions: where search is placed, what candidates are, whether feedback represents trustworthiness. Trend extends from pre-deployment search for a single robust solution to adaptive protection during training and execution.
Trustworthy Self-Evolving AI
Constrained Exploration of Search Space
First mechanism answers "what updates are allowed". MermaidFlow evolves typed workflows under structural/semantic constraints, excluding illegal, non-executable, incompatible offspring. SEVerA combines iterative synthesis with formal verification so candidates satisfying explicit behavioral properties enter subsequent learning. This constrains feasible update space. Structural correctness does not guarantee full semantic safety; formal verification only covers specified properties; more explicit constraints make guarantees clearer.
Selection-Guided Improvement
Second mechanism allows multiple modifications but uses safety-aware fitness, verification, or admission rules to decide retention. AgentBreeder jointly selects multi-agent scaffolds for capability and safety; HarnessBank separates candidate modifications (prompts, knowledge, tools, runtime controls) from long-term archival decisions. SHE converts trajectory-level failures into local safety updates, checking safety and utility; immediate memory systems use test-verify-write loops to avoid unverified experience persisting. TAME, SkillHarness, Membrane further govern memory, skill boundaries, continuous reuse of safety knowledge, while addressing jailbreak and over-refusal. These methods manage when inherited state applies and when it needs correction. Verification mechanisms have coverage boundaries, not universal safety guarantees for arbitrary future environments.
Strategic Attack-Defense Co-Evolution
Third mechanism lets new attacks change defense training, defense improvements change attack selection pressure. ROMANCE maintains diverse attackers and alternates training cooperative policies; CEMMA evolves multimodal attacks then uses successful cases to update defenders; Evo-MARL combines attack prompt populations with multi-agent RL. More persistent systems preserve both sides' states: DARWIN maintains attack strategy pool and continuous defense versions; EvoSafety externalizes attack skills and verified defense memories. Both sides need not use same optimization algorithm; key is mutual feedback change and cross-round knowledge accumulation.
Discussion
Self-evolution advances robustness from a static version to the entire update process. Erroneous memory, unreliable skills, fragile processes once retained can affect later tasks and modifications. The three mechanisms govern feasible space, retention decisions, adaptive feedback respectively; they complement each other, not three mutually exclusive alternatives. Passing current version tests does not substitute for trustworthy evaluation of the update process. As populations, archives, external storage continuously accumulate, selection and inheritance themselves become system behaviors needing scrutiny.
Open Challenges and Future Research Directions
Multi-Component Search Under Limited Budget
Agents contain prompts, retrieval, memory, tools, workflows, model configurations; joint evolution expands search space and evaluation cost. Different modifications have different costs: prompts/routing may need few queries; persistent memory/tool changes need broader execution verification; parameter updates involve training compute. Authors suggest prioritizing low-cost, high-impact components with budget-aware allocation. Multi-agent settings must consider coupling risks between components and agents. This is a future design direction, not a validated universal optimal scheduling scheme.
Multi-Solution Defense for Diverse Threats
Single architecture/policy/protection may not cover different attacks, data distributions, safety requirements. Evolutionary search can retain multiple complementary solutions deployed via ensemble, voting, contextual selection. Need genuine defense complementarity, not many behaviorally similar candidates; runtime cost and selection mechanisms must be evaluated.
Multi-Task Generalization in Self-Evolution
Agents work across communication, retrieval, planning, coding, tool use, but updates often evaluated only on few tasks that produced them. Local improvements may not transfer, even harm existing capabilities. Authors suggest cross-task candidate evaluation, selecting updates that generalize while retaining old capabilities, linking stability-plasticity trade-off in continual learning.
Multi-Objective Benchmarks for Trustworthiness and Evolutionary Search
Current safety benchmarks and optimization benchmarks are often separate, making it hard to judge if same safety gain consumes same resources. Future unified benchmarks should jointly report robustness, safety, task utility, search cost, convergence under comparable budgets. Weighted aggregate scores can assist summarization but must not obscure individual metrics and their trade-offs, nor replace behavioral coverage analysis.
Conclusion
The survey's contribution is a unified organizational perspective, not a single safety algorithm winning in all scenarios. It connects attack, defense, persistent self-evolution around "what evolves, how to search, how to evaluate": attacks use populations to find failures, defenses search data/structure/strategy, self-evolution further governs which changes can be inherited. Evolutionary computation's strengths: adapts to black-box and structured search, explicitly preserves multi-objective trade-offs, maintains behaviorally diverse candidates. But trustworthiness does not automatically appear from using evolutionary mechanisms: wrong fitness, narrow test scope, expensive/noisy feedback still yield unreliable conclusions. For trustworthy AI research and engineering, the key reminder: evaluate both results and search process; consider both current utility and future inheritance. Especially when systems continuously write memory, skills, defense archives, "generate-screen-verify-retain" is not just an optimization flow but a governance boundary of trustworthy systems.
Code example
来源:专知
本文
约7400字
,建议阅读
10
分钟
这篇综述以进化计算为共同视角,串联三条研究线。Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Party THU
Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
