The Rise of Agentic AI: A Review of Definitions, Frameworks, Architectures, Applications, Evaluation Metrics, and Challenges
đŻ Resumo
Abstract Original
Agentic AI systems are a recently emerged and important approach that goes beyond traditional AI, generative AI, and autonomous systems by focusing on autonomy, adaptability, and goal-driven reasoning. This study provides a clear review of agentic AI systems by bringing together their definitions, frameworks, and architectures, and by comparing them with related areas like generative AI, autonomic computing, and multi-agent systems. To do this, we reviewed 143 primary studies on current LLM-based and non-LLM-driven agentic systems and examined how they support planning, memory, reflection, and goal pursuit. Furthermore, we classified architectural models, inputâoutput mechanisms, and applications based on their task domains where agentic AI is applied, supported using tabular summaries that highlight real-world case studies. Evaluation metrics were classified as qualitative and quantitative measures, along with available testing methods of agentic AI systems to check the systemâs performance and reliability. This study also highlights the main challenges and limitations of agentic AI, covering technical, architectural, coordination, ethical, and security issues. We organized the conceptual foundations, available tools, architectures, and evaluation metrics in this research, which defines a structured foundation for understanding and advancing agentic AI. These findings aim to help researchers and developers build better, clearer, and more adaptable systems that support responsible deployment in different domains.
đ Bibliografia
BANDI, Ajay et al. The Rise of Agentic AI: A Review of Definitions, Frameworks, Architectures, Applications, Evaluation Metrics, and Challenges. Future Internet, v. 17, n. 9, p. 404, 4 set. 2025.
đ§ Minhas Notas & AnĂĄlise
đŻ Objetivo
- Fornecer uma revisĂŁo abrangente da IA agĂȘntica, sintetizando suas definiçÔes, frameworks e arquiteturas.
- Comparar a IA agĂȘntica com paradigmas relacionados, como a InteligĂȘncia Artificial Generativa, a computação autonĂŽmica e os Sistemas Multiagentes.
- Identificar as ferramentas e frameworks atuais que suportam capacidades agĂȘnticas.
- Descrever os modelos arquitetĂŽnicos disponĂveis e explorar suas aplicaçÔes em diversos domĂnios.
- Analisar os mecanismos de entrada e saĂda, os mĂ©todos de teste existentes e as mĂ©tricas usadas para avaliar o desempenho da IA agĂȘntica.
- Abordar os principais desafios e limitaçÔes da ĂĄrea para estabelecer uma base estruturada para a compreensĂŁo e o avanço da IA agĂȘntica.
𧏠Método
- Realização de uma revisão de literatura que extraiu dados de 143 estudos primårios publicados entre o ano de 2005 e o presente.
- Busca executada em bases de dados acadĂȘmicas importantes, incluindo Library Search, ACM Digital Library, ScienceDirect, Google Scholar e Semantic Scholar.
- Criação de uma string de busca precisa combinando termos centrais como âagentic AIâ, âmultiagent systemsâ, âautonomic computingâ e palavras-chave relacionadas a revisĂ”es e levantamentos bibliogrĂĄficos.
- InclusĂŁo de artigos que discutiam conceitos, sistemas ou desafios da IA agĂȘntica atravĂ©s de estudos de caso, pesquisas e outros mĂ©todos.
- AnĂĄlise adicional de 21 artigos de revisĂŁo e levantamentos para fornecer perspectivas complementares sobre a IA agĂȘntica e preencher lacunas de dimensĂ”es menos desenvolvidas na literatura anterior
đ Resultados
- A IA agĂȘntica foi definida conceitualmente como sistemas autĂŽnomos orientados a objetivos que operam por longos perĂodos com mĂnima supervisĂŁo humana, distinguindo-se da IA tradicional por sua capacidade de planejamento e raciocĂnio contextual.
- Identificação de seis componentes principais que formam a arquitetura dos sistemas agĂȘnticos: percepção e modelagem de mundo, memĂłria, planejamento e raciocĂnio, execução e atuação, reflexĂŁo e avaliação, e orquestração.
- Classificação de cinco modelos arquitetĂŽnicos de orquestração: agente Ășnico ReAct, supervisor/hierĂĄrquico, hĂbrido reativo-deliberativo, arquitetura BDI (crença-desejo-intenção) e decisĂŁo em camadas (neuro-simbĂłlica).
- Mapeamento de aplicaçÔes prĂĄticas da IA agĂȘntica em mĂșltiplas indĂșstrias, como saĂșde, militar, transporte, desenvolvimento de software, finanças e cidades inteligentes.
- Classificação das métricas de avaliação em medidas qualitativas (ex: Explicabilidade, satisfação do usuårio, mitigação de viés) e medidas quantitativas (ex: Precisão, Tempo de Conclusão da Tarefa, taxa de cliques).
- Destaque para os principais desafios atuais da årea, que englobam limitaçÔes técnicas/arquitetÎnicas, problemas de coordenação multiagente, questÔes de interação humano-IA e riscos éticos e de segurança.
đ ComentĂĄrios Extras
- O artigo deixa claro que a evolução para a IA agĂȘntica representa uma mudança de sistemas reativos que respondem a prompts individuais para colaboradores proativos e autĂŽnomos.
- A combinação de Modelos de Linguagem Grande (LLMs) com Aprendizado por Reforço (RL) Ă© apontada como a base forte que permite aos agentes interagir e melhorar suas decisĂ”es continuamente atravĂ©s da experiĂȘncia.
- Apesar dos avanços, a adoção corporativa e no mundo real exige a superação de gargalos importantes, especialmente no que diz respeito Ă segurança contĂnua, governança e TransparĂȘncia das decisĂ”es tomadas pelos agentes (efeito caixa-preta).
đ ConexĂ”es do Cofre
- Temas Relacionados: Agentes AutÎnomos, Grandes modelos de linguagem, Aprendizado por Reforço (RL), Sistemas Multiagentes (MAS), IA Generativa, Arquitetura de Software, Teste e Avaliação de IA.
- Conceitos & Ferramentas Detalhados:
- InteligĂȘncia Artificial AgĂȘntica
- Sistemas Multiagentes
- Métricas de Avaliação de Modelos
- Matriz de ConfusĂŁo
- Acuråcia | Precisão | Revocação | F1 Score
- Distùncia de Edição de Grafo (GED) | Tempo de Conclusão da Tarefa (TCT)
- Robustez | Adaptabilidade | TransparĂȘncia | Explicabilidade
- LangChain | AutoGPT | BabyAGI | AutoGen | SuperAGI | OpenAgents
- Communicative Agents for Mind Exploration or Lasge-Scale Language Model Society (CAMEL)
- Task-Based Cognitive Sequential Planning Network (TB-CSPN)
- Trabalhos que citam este:
đŁ Bibliografia Latex
@article{bandi2025,
title = {The {{Rise}} of {{Agentic AI}}: {{A Review}} of {{Definitions}}, {{Frameworks}}, {{Architectures}}, {{Applications}}, {{Evaluation Metrics}}, and {{Challenges}}},
shorttitle = {The {{Rise}} of {{Agentic AI}}},
author = {Bandi, Ajay and Kongari, Bhavani and Naguru, Roshini and Pasnoor, Sahitya and Vilipala, Sri Vidya},
year = 2025,
month = 9,
journal = {Future Internet},
volume = {17},
number = {9},
pages = {404},
issn = {1999-5903},
doi = {10.3390/fi17090404},
urldate = {2026-06-15},đïž Notas e Destaques
- This study provides a clear review of agentic AI systems by bringing together their definitions, frameworks, and architectures, and by comparing them with related areas like generative AI, autonomic computing, and multi-agent systems. (p. 1)
- To do this, we reviewed 143 primary studies on current LLM-based and non-LLM-driven agentic systems and examined how they support planning, memory, reflection, and goal pursuit. (p. 1)
- Furthermore, we classified architectural models, inputâoutput mechanisms, and applications based on their task domains where agentic AI is applied, supported using tabular summaries that highlight real-world case studies (p. 1)
- Evaluation metrics were classified as qualitative and quantitative measures, along with available testing methods of agentic AI systems to check the systemâs performance and reliability. (p. 1)
- Agentic AI refers to AI systems that do not just answer prompts; they set sub-goals, choose tools, and take multi-step actions to achieve a userâs objective with limited supervision. (p. 1)
- An agent is a system that senses its environment, decides what to do, and takes action to achieve a goal. (p. 1)
- The term AI agent refers to an autonomous software entity that perceives, reasons, and acts to complete specific adaptively directed tasks. (p. 2)
- The term agentic AI refers to a multi-agent system where specialized agents collaborate, coordinate, and plan to achieve complex, high-level objectives. (p. 2)
- This review synthesizes how the field defines agentic AI, the architectures that enable it, where it is being applied, how it is measured, and the open challenges, such as reliability, coordination, safety, and governance, that must be addressed for widespread deployment. (p. 2)
- The purpose of this research is to provide a comprehensive review of agentic AI by synthesizing its definitions, frameworks, and architectures, while comparing agentic AI from related paradigms. This study aims to find current tools and frameworks that support agentic capabilities, describe available architectural models of such systems, and explore their applications across domains. Furthermore, it concentrates on the inputâoutput mechanism of agentic AI, existing testing methods, and metrics used to assess agentic performance, and addresses the key challenges and limitations. By doing so, the paper establishes a structured foundation for understanding, assessing, and advancing the field of agentic AI. (p. 3)
- To carry out this study, we searched several academic databases, including Library Search, ACM Digital Library, ScienceDirect, Google Scholar, and Semantic Scholar. We created a precise search query (âagentic AIâ OR âagentâ OR âautonomous AIâ OR âmultiagent systemsâ OR âautonomic computingâ OR âgenerative AIâ OR âAI planningâ OR âAI memory systemsâ OR âAI feedback loopsâ) AND (âsurveyâ OR âoverviewâ OR âreviewâ OR âsummaryâ OR âliterature reviewâ) to find studies relevant to our research. (p. 5)
- We included papers that discussed agentic AI concepts, systems, or challenges, using case studies, surveys, and other research methods. (p. 5)
- we used these 143 primary studies to extract the data, which were published from the year 2005 to the present (p. 5)
- 32 of these selected papers are published on the arXiv platform (p. 5) ComentĂĄrio: Alto nĂșmero de publicaçÔes que nĂŁo passaram pelo processo de revisĂŁo por pares e que provavelmente contĂ©m problemas.
- We observed that most of the publications are multidisciplinary, which makes sense in the context of agentic AI, as these disciplines are focusing on solving their problems using agentic AI. (p. 5)
- From this analysis, it is clear that most existing surveys cover only selected dimensions, leaving important areas such as standardized benchmarks, performance evaluation, and inputâoutput modeling less developed. (p. 6)
- Agentic artificial intelligence (agentic AI) marks a major change in how AI systems are designed. Moving beyond passive and reactive tools, agentic AI refers to autonomous, goaldriven systems that can operate on their own for long periods, requiring minimal human supervision [27,31â33]. Unlike traditional AI or standard large language models (LLMs) that respond only to single prompts [34], these systems can understand broad objectives, break them down into smaller tasks, and carry out multi-step plans while adapting to feedback from changing environments [35â37]. Their âagenticâ nature also highlights their ability to take on responsibilities from humans [32], act purposefully, and be accountable for the results they produce. (p. 7)
- What makes agentic AI stand out is its combination of key abilities in one system. These include strategic planning, memory that preserves context over time, using external tools to extend capabilities, and collaborating with other agents [38â40]. (p. 8)
- At the heart of agentic AI is self-directed, goal-focused intelligence. These systems operate with a âdegree of agenticness,â meaning they can work proactively, plan ahead, and shape outcomes over time [14,46]. (p. 8)
- Architecturally, they often integrate large AI models with knowledge bases, planners, access to external tools, and persistent memory to improve understanding and independence [21,42,52]. (p. 8)
- Agentic AI is built strongly on top of LLMs and RL. (p. 9)
- The development of LLM-based frameworks like LangChain [26,31,78,80], AutoGPT [26,31,78], BabyAGI [31], AutoGen [80], and OpenAgents [26] shows a clear trend toward building smarter, more independent AI systems. These frameworks give AI agents the ability to plan, remember past actions, reflect on their performance, pursue goals, and use external tools. By equipping agents with these cognitive skills, they become more adaptable and capable of handling complex tasks with minimal human guidance. (p. 10)
- LangChain is a popularly used open source framework that connects LLM to external tools, data sources and APIs, which simplifies the work of constructing complex workflows with multiple steps [26,31,78,79]. (p. 10)
- AutoGPT is built as a fully autonomous agent capable of completing complex tasks without needing constant human guidance [43,45,47]. (p. 10)
- Intelligent agents can act and make decisions outside of large language models. NonLLM-dependent agentic systems achieve autonomy through either rules, learning, planning, or reactive behaviors; they have been used extensively in robotics, simulations, and automated decision-making. (p. 12)
- One important type is rule-based multi-agent systems. In these systems, agents follow predefined rules or logic to interact and achieve coordinated outcomes (p. 12)
- The model, typically an LLM, large agent model (LAM), or foundation model (FM), provides the systemâs core intelligence. It often serves as the reasoning engine, perceptual front-end, and, in many cases, the orchestrator. In Multi-Agent Artificial Intelligence (MAAI) frameworks, the model forms a base layer for perception, action, orchestration, workflows, and user interaction [96]. (p. 13)
- A common integration pattern is a layered flow that couples components into an end-to-end decision loop: perception â world/state modeling â planning/reasoning â execution/tooling â evaluation/feedback â memory â (back to) planning (p. 20)
- Centralized orchestration: Many systems elevate an LLM to a supervisor role that maintains shared context, routes tasks to specialist modules/agents (planner, retriever, executor, critic), triggers reflection or replanning, and arbitrates conflicts among goals (p. 20)
- Regardless of centralization, integrating components with an explicit workflow graph improves robustness under non-determinism. Nodes encapsulate component calls (with contracts for inputs/outputs and budgets), and edges encode control logic (success/failure, timeouts, escalation). (p. 20)
- Agentic AI systems are used to complete routine activities, support in real-time decision-making, handle complex tasks or problems, and provide personalized assistance by adapting to the context. The systemâs strength is in combining data from many sources, making ethical and informed decisions, and coordinating multiple agents to work together in real time. The rationale for creating a domain-based categorization of agentic AIs is to clarify how these systems are molded and shaped by the specific needs, boundaries, and objectives of each of the fields deployed. (p. 22)
- Explainability: The degree to which the system produces clear and understandable reasons or explanations for its actual decisions or outputs. (p. 36)
- Transparency: A measure of how openly the AI systemâs inner workings and decision processes are made visible or available for inspection (p. 36)
- User Satisfaction: It is a measure of how well the agentic AI system meets the needs, expectations, and preferences of users [163]. (p. 36)
- Fairness: It can be used to identify biases system in agentic AI, for example, responding in different ways, or differing response quality in different demographic groups [56] (p. 36)
- Bias Mitigation: It can be used to reduce or eliminate bias from agentic AI systems. If agentic AI systems that amplify societal biases can lead to unfair outcomes, [49], particularly oscillating more aggregated biases in regards to vulnerable groups. (p. 36)
- Co-operative Behavior: It can be used to measure how well AI agents collaborate and coordinate with one another to achieve mutual goals. (p. 36)
- Adaptability: It encompasses ongoing interpersonal learning and agile feedback and tuned goals and measures via conceptual multi-agent models, formal policy operationalizations such as Petri Nets [154], and applied second burst benchmark measures that simulate unpredictable, complex, and dynamic environments [75]. (p. 36)
- Robustness: It is a measure of the ability of an agentic AI system to maintain goaldirected performance [48] despite internal failures and external adversities. (p. 36)
- Precision: It is a measure of how many were actually correct out of all positive predictions made by a model, to find the accuracy of positive predictions. (p. 37)
- Recall: It measures how many the model gets right out of all actual positives, to obtain all positive instances. (p. 37)
- F1 Score: It is a measure that combines precision and recall into a single number by calculating their harmonic mean. (p. 37)
- Graph Edit Distance (GED): An important structural metric in agentic AI, using the level of structural similarity between an AI-generated task graph and a ground-truth graph as a measurement. (p. 37)
- Rule Fidelity: A measure of how accurately the symbolic, human-readable rules generated by the system reflect the actual decision-making process. (p. 37)
- Task Completion Time (TCT): It measures the time taken by the agent to plan, execute, and complete a particular task. (p. 37)
- Click-Through Rate (CTR): It assesses how good of a job the agentic AI system interacts with the user to click a recommendation or content we list to them, e.g, associating CTR with the systemâs relevance, impact, etc. Models like the GPT-4 (ColdLLM)6, the CG4CTR6 and the LLM-InS6 simulators were used to test if it was possible to associate different scenarios of CTR to the recommendation list through simulations of common interest content personalization, visual design impact, Inter environmental dialog personality, or deferred dialog. (p. 37)
- Gross Merchandise Value (GMV): It measures the total dollar amount of the items sold or the total dollar amount of completed transactions due to the actions of, or recommendations made by, the AI. (p. 37)