En Parte 1 De esta serie tutorial, presentamos Agentes de IAprogramas autónomos que realizan tareas, toman decisiones y se comunican con otros.
En Parte 2 De esta serie tutorial, entendimos cómo hacer que el agente intente volver a intentar hasta que la tarea se complete a través de Iteraciones y cadenas.
Un solo agente generalmente puede funcionar de manera efectiva utilizando una herramienta, pero puede ser menos efectivo cuando se usa muchas herramientas simultáneamente. Una forma de abordar tareas complicadas es a través de un enfoque de “dividir y concebir”: crear un agente especializado para cada tarea y hacer que trabajen juntos como un Sistema de múltiples agentes (MAS).
En un MAS, múltiples agentes colaboran para lograr objetivos comunes, a menudo abordando desafíos que son demasiado difíciles de manejar para un solo agente solo. Hay dos formas principales en que pueden interactuar:
- Flujo secuencial – Los agentes hacen su trabajo en un orden específico, uno tras otro. Por ejemplo, el Agente 1 termina su tarea, y luego el Agente 2 usa el resultado para hacer su tarea. Esto es útil cuando las tareas dependen entre sí y deben hacerse paso a paso.
- Flujo jerárquico – Por lo general, un agente de nivel superior administra todo el proceso y proporciona instrucciones a los agentes de nivel más bajo que se centran en tareas específicas. Esto es útil cuando la salida final requiere algo de ida y vuelta.
En este tutorial, voy a mostrar cómo construir desde cero diferentes tipos de sistemas de múltiples agentesde simple a más avanzado. Presentaré un código de Python útil que se puede aplicar fácilmente en otros casos similares (solo copie, pegue, ejecute) y camine por cada línea de código con comentarios para que pueda replicar este ejemplo (enlace al código completo al final del artículo).
Configuración
Consulte Parte 1 para la configuración de Ollama y el principal LLM.
import ollama
llm = "qwen2.5"
En este ejemplo, le pediré al modelo que procese imágenes, por lo tanto, también voy a necesitar un Visión LLM. Es una versión especializada de un modelo de lenguaje grande que, integrando NLP con CV, está diseñada para comprender las entradas visuales, como imágenes y videos, además del texto.
Microsoft’s Llava es una opción eficiente, ya que también puede funcionar sin una GPU.
Después de completar la descarga, puede pasar a Python y comenzar a escribir código. Cargamos una imagen para que podamos probar el Vision LLM.
from matplotlib import image as pltimg, pyplot as plt
image_file = "draghi.jpeg"
plt.imshow(pltimg.imread(image_file))
plt.show()
Para probar el Vision LLM, puede pasar la imagen como entrada:
import ollama
ollama.generate(model="llava",
prompt="describe the image",
images=[image_file])["response"]
Secuencial Sistema de múltiples agentes
Construiré dos agentes que trabajen en un flujo secuencialuno tras otro, donde el segundo toma la salida de la primera como entrada, como una cadena.
- El primer agente debe procesar una imagen proporcionada por el usuario y devolver una descripción verbal de lo que ve.
- El segundo agente Buscará Internet e intentará comprender dónde y cuándo se tomó la imagen, en función de la descripción proporcionada por el primer agente.
Ambos agentes usarán uno Herramienta cada. El primer agente tendrá el Vision LLM como herramienta. Por favor recuerde eso con Ollamapara usar una herramienta, la función debe describirse en un diccionario.
def process_image(path: str) -> str:
return ollama.generate(model="llava", prompt="describe the image", images=[path])["response"]
tool_process_image = {'type':'function', 'function':{
'name': 'process_image',
'description': 'Load an image for a given path and describe what you see',
'parameters': {'type': 'object',
'required': ['path'],
'properties': {
'path': {'type':'str', 'description':'the path of the image'},
}}}}
El segundo agente debe tener una herramienta de búsqueda en la web. En los artículos anteriores de esta serie tutorial, mostré cómo aprovechar el Duckduckgo Paquete para buscar en la web. Entonces, esta vez, podemos usar una nueva herramienta: Wikipedia (pip install wikipedia==1.4.0). Puede usar directamente la biblioteca original o importar el Langchain envoltura.
from langchain_community.tools import WikipediaQueryRun
from langchain_community.utilities import WikipediaAPIWrapper
def search_wikipedia(query:str) -> str:
return WikipediaQueryRun(api_wrapper=WikipediaAPIWrapper()).run(query)
tool_search_wikipedia = {'type':'function', 'function':{
'name': 'search_wikipedia',
'description': 'Search on Wikipedia by passing some keywords',
'parameters': {'type': 'object',
'required': ['query'],
'properties': {
'query': {'type':'str', 'description':'The input must be short keywords, not a long text'},
}}}}
## test
search_wikipedia(query="draghi")
Primero, debe escribir un mensaje para describir la tarea de cada agente (cuanto más detallado, mejor), y ese será el primer mensaje en el historial de chat con el LLM.
prompt = '''
You are a photographer that analyzes and describes images in details.
'''
messages_1 = [{"role":"system", "content":prompt}]
Una decisión importante de tomar al construir un MAS es si los agentes deben compartir el historial de chat o no. El Gestión de la historia del chat Depende del diseño y los objetivos del sistema:
- Historial de chat compartido – Los agentes tienen acceso a un registro de conversación común, lo que les permite ver lo que otros agentes han dicho o hecho en interacciones anteriores. Esto puede mejorar la colaboración y la comprensión del contexto general.
- Historial de chat separado – Los agentes solo tienen acceso a sus propias interacciones, centrándose solo en su propia comunicación. Este diseño se usa típicamente cuando la toma de decisiones independientes es importante.
Recomiendo mantener los chats separados a menos que sea necesario hacer lo contrario. Los LLM pueden tener una ventana de contexto limitada, por lo que es mejor hacer que la historia sea lo más lite posible.
prompt = '''
You are a detective. You read the image description provided by the photographer, and you search Wikipedia to understand when and where the picture was taken.
'''
messages_2 = [{"role":"system", "content":prompt}]
Por conveniencia, utilizaré la función definida en los artículos anteriores para procesar la respuesta del modelo.
def use_tool(agent_res:dict, dic_tools:dict) -> dict:
## use tool
if "tool_calls" in agent_res["message"].keys():
for tool in agent_res["message"]["tool_calls"]:
t_name, t_inputs = tool["function"]["name"], tool["function"]["arguments"]
if f := dic_tools.get(t_name):
### calling tool
print('🔧 >', f"x1b[1;31m{t_name} -> Inputs: {t_inputs}x1b[0m")
### tool output
t_output = f(**tool["function"]["arguments"])
print(t_output)
### final res
res = t_output
else:
print('🤬 >', f"x1b[1;31m{t_name} -> NotFoundx1b[0m")
## don't use tool
if agent_res['message']['content'] != '':
res = agent_res["message"]["content"]
t_name, t_inputs = '', ''
return {'res':res, 'tool_used':t_name, 'inputs_used':t_inputs}
Como ya hicimos en tutoriales anteriores, la interacción con los agentes puede iniciarse con un Mientras que el bucle. Se solicita al usuario que proporcione una imagen que procesará el primer agente.
dic_tools = {'process_image':process_image,
'search_wikipedia':search_wikipedia}
while True:
## user input
try:
q = input('📷 > give me the image to analyze:')
except EOFError:
break
if q == "quit":
break
if q.strip() == "":
continue
messages_1.append( {"role":"user", "content":q} )
plt.imshow(pltimg.imread(q))
plt.show()
## Agent 1
agent_res = ollama.chat(model=llm,
tools=[tool_process_image],
messages=messages_1)
dic_res = use_tool(agent_res, dic_tools)
res, tool_used, inputs_used = dic_res["res"], dic_res["tool_used"], dic_res["inputs_used"]
print("👽📷 >", f"x1b[1;30m{res}x1b[0m")
messages_1.append( {"role":"assistant", "content":res} )
El primer agente utilizó la herramienta Vision LLM y el texto reconocido dentro de la imagen. Ahora, la descripción se pasará al segundo agente, que extraerá algunas palabras clave para buscar Wikipedia.
## Agent 2
messages_2.append( {"role":"system", "content":"-Picture: "+res} )
agent_res = ollama.chat(model=llm,
tools=[tool_search_wikipedia],
messages=messages_2)
dic_res = use_tool(agent_res, dic_tools)
res, tool_used, inputs_used = dic_res["res"], dic_res["tool_used"], dic_res["inputs_used"]
El segundo agente utilizó la herramienta y extraía información de la web, según la descripción proporcionada por el primer agente. Ahora, puede procesar todo y dar una respuesta final.
if tool_used == "Search_wikipedia": Messages_2.Append ({"rol": "System", "Content": "-Wikipedia:"+Res}) agent_res = ollama.chat (model = llm, herramientas =[]Messages = Messages_2) Dic_res = Use_Tool (Agent_Res, Dic_Tools) Res, Tool_used, Inputs_Used = DIC_RES["res"]Dic_Res["tool_used"]Dic_Res["inputs_used"]
else: Messages_2.append ({"rol": "Asistente", "Contenido": Res}) Impresa ("👽📖>", f "x1b[1;30m{res}x1b[0m")
This is literally perfect! Let’s move on to the next example.
Hierarchical Multi-Agent System
Imagine having a squad of Agents that operates with a hierarchical flow, just like a human team, with distinct roles to ensure smooth collaboration and efficient problem-solving. At the top, a manager oversees the overall strategy, talking to the customer (the user), making high-level decisions, and guiding the team toward the goal. Meanwhile, other team members handle operative tasks. Just like humans, Agents can work together and delegate tasks appropriately.
I shall build a tech team of 3 Agents with the objective of querying a SQL database per user’s request. They must work in a hierarchical flow:
- The Lead Agent talks to the user and understands the request. Then, it decides which team member is the most appropriate for the task.
- The Junior Agent has the job of exploring the db and building SQL queries.
- The Senior Agent shall review the SQL code, correct it if necessary, and execute it.
LLMs know how to code by being exposed to a large corpus of both code and natural language text, where they learn patterns, syntax, and semantics of programming languages. The model learns the relationships between different parts of the code by predicting the next token in a sequence. In short, LLMs can generate SQL code but can’t execute it, Agents can.
First of all, I am going to create a database and connect to it, then I shall prepare a series of Tools to execute SQL code.
## Read dataset
import pandas as pd
dtf = pd.read_csv('http://bit.ly/kaggletrain')
dtf.head(3)
## Create dbimport sqlite3
dtf.to_sql(index=False, name="titanic",
con=sqlite3.connect("database.db"),
if_exists="replace")
## Connect db
from langchain_community.utilities.sql_database import SQLDatabase
db = SQLDatabase.from_uri("sqlite:///database.db")
Let’s start with the Junior Agent. LLMs don’t need Tools to generate SQL code, but the Agent doesn’t know the table names and structure. Therefore, we need to provide Tools to investigate the database.
from langchain_community.tools.sql_database.tool import ListSQLDatabaseTool
def get_tables() -> str:
return ListSQLDatabaseTool(db=db).invoke("")
tool_get_tables = {'type':'function', 'function':{
'name': 'get_tables',
'description': 'Returns the name of the tables in the database.',
'parameters': {'type': 'object',
'required': []'Propiedades': {}}}} ## test get_tables ()
Eso mostrará las tablas disponibles en el DB, y esto imprimirá las columnas en una tabla.
from langchain_community.tools.sql_database.tool import InfoSQLDatabaseTool
def get_schema(tables: str) -> str:
tool = InfoSQLDatabaseTool(db=db)
return tool.invoke(tables)
tool_get_schema = {'type':'function', 'function':{
'name': 'get_schema',
'description': 'Returns the name of the columns in the table.',
'parameters': {'type': 'object',
'required': ['tables'],
'properties': {'tables': {'type':'str', 'description':'table name. Example Input: table1, table2, table3'}}
}}}
## test
get_schema(tables='titanic')
Dado que este agente debe usar más de una herramienta que pueda fallar, escribiré un aviso sólido, siguiendo la estructura del artículo anterior.
prompt_junior = '''
[GOAL] You are a data engineer who builds efficient SQL queries to get data from the database.
[RETURN] You must return a final SQL query based on user's instructions.
[WARNINGS] Use your tools only once.
[CONTEXT] In order to generate the perfect SQL query, you need to know the name of the table and the schema.
First ALWAYS use the tool 'get_tables' to find the name of the table.
Then, you MUST use the tool 'get_schema' to get the columns in the table.
Finally, based on the information you got, generate an SQL query to answer user question.
'''
Moviéndose hacia el Agente superior. La verificación de código no requiere ningún truco en particular, solo puede usar el LLM.
def sql_check(sql: str) -> str:
p = f'''Double check if the SQL query is correct: {sql}. You MUST just SQL code without comments'''
res = ollama.generate(model=llm, prompt=p)["response"]
return res.replace('sql','').replace('```','').replace('n',' ').strip()
tool_sql_check = {'type':'function', 'function':{
'name': 'sql_check',
'description': 'Before executing a query, always review the SQL query and correct the code if necessary',
'parameters': {'type': 'object',
'required': ['sql'],
'properties': {'sql': {'type':'str', 'description':'SQL code'}}
}}}
## test
sql_check(sql='SELECT * FROM titanic TOP 3')
La ejecución del código en la base de datos es una historia diferente: LLMS no puede hacerlo solo.
from langchain_community.tools.sql_database.tool import QuerySQLDataBaseTool
def sql_exec(sql: str) -> str:
return QuerySQLDataBaseTool(db=db).invoke(sql)
tool_sql_exec = {'type':'function', 'function':{
'name': 'sql_exec',
'description': 'Execute a SQL query',
'parameters': {'type': 'object',
'required': ['sql'],
'properties': {'sql': {'type':'str', 'description':'SQL code'}}
}}}
## test
sql_exec(sql='SELECT * FROM titanic LIMIT 3')
Y, por supuesto, un buen aviso.
prompt_senior = '''[GOAL] You are a senior data engineer who reviews and execute the SQL queries written by others.
[RETURN] You must return data from the database.
[WARNINGS] Use your tools only once.
[CONTEXT] ALWAYS check the SQL code before executing on the database.First ALWAYS use the tool 'sql_check' to review the query. The output of this tool is the correct SQL query.You MUST use ONLY the correct SQL query when you use the tool 'sql_exec'.'''
Finalmente, crearemos el Agente principal. Tiene el trabajo más importante: invocar a otros agentes y decirles qué hacer. Hay muchas maneras de lograr eso, pero encuentro que crear una herramienta simple la más precisa.
def invoke_agent(agent:str, instructions:str) -> str:
return agent+" - "+instructions if agent in ['junior','senior'] else f"Agent '{agent}' Not Found"
tool_invoke_agent = {'type':'function', 'function':{
'name': 'invoke_agent',
'description': 'Invoke another Agent to work for you.',
'parameters': {'type': 'object',
'required': ['agent', 'instructions'],
'properties': {
'agent': {'type':'str', 'description':'the Agent name, one of "junior" or "senior".'},
'instructions': {'type':'str', 'description':'detailed instructions for the Agent.'}
}
}}}
## test
invoke_agent(agent="intern", instructions="build a query")
Describa en el mensaje qué tipo de comportamiento esperas. Trate de ser lo más detallado posible, ya que los sistemas jerárquicos de múltiples agentes pueden ser muy confusos.
prompt_lead = '''
[GOAL] You are a tech lead.
You have a team with one junior data engineer called 'junior', and one senior data engineer called 'senior'.
[RETURN] You must return data from the database based on user's requests.
[WARNINGS] You are the only one that talks to the user and gets the requests from the user.
The 'junior' data engineer only builds queries.
The 'senior' data engineer checks the queries and execute them.
[CONTEXT] First ALWAYS ask the users what they want.
Then, you MUST use the tool 'invoke_agent' to pass the instructions to the 'junior' for building the query.
Finally, you MUST use the tool 'invoke_agent' to pass the instructions to the 'senior' for retrieving the data from the database.
'''
Mantendré el historial de chat separado para que cada agente conozca solo una parte específica de todo el proceso.
dic_tools = {'get_tables':get_tables,
'get_schema':get_schema,
'sql_exec':sql_exec,
'sql_check':sql_check,
'Invoke_agent':invoke_agent}
messages_junior = [{"role":"system", "content":prompt_junior}]
messages_senior = [{"role":"system", "content":prompt_senior}]
messages_lead = [{"role":"system", "content":prompt_lead}]
Todo está listo para Comience el flujo de trabajo. Después de que el usuario comienza el chat, el primero en responder es el líder, que es el único que interactúa directamente con el humano.
while True:
## user input
q = input('🙂 >')
if q == "quit":
break
messages_lead.append( {"role":"user", "content":q} )
## Lead Agent
agent_res = ollama.chat(model=llm, messages=messages_lead, tools=[tool_invoke_agent])
dic_res = use_tool(agent_res, dic_tools)
res, tool_used, inputs_used = dic_res["res"], dic_res["tool_used"], dic_res["inputs_used"]
agent_invoked = res.split("-")[0].strip() if len(res.split("-")) > 1 else ''
instructions = res.split("-")[1].strip() if len(res.split("-")) > 1 else ''
###-->CODE TO INVOKE OTHER AGENTS HERE<--###
## Lead Agent final response print("👩💼 >", f"x1b[1;30m{res}x1b[0m") messages_lead.append( {"role":"assistant", "content":res} )
El agente principal decidió invocar al agente junior dándole algunas instrucciones, según la interacción con el usuario. Ahora el agente junior comenzará a trabajar en la consulta.
## Invoke Junior Agent
if agent_invoked == "junior":
print("😎 >", f"x1b[1;32mReceived instructions: {instructions}x1b[0m")
messages_junior.append( {"role":"user", "content":instructions} )
### use the tools
available_tools = {"get_tables":tool_get_tables, "get_schema":tool_get_schema}
context = ''
while available_tools:
agent_res = ollama.chat(model=llm, messages=messages_junior,
tools=[v for v in available_tools.values()])
dic_res = use_tool(agent_res, dic_tools)
res, tool_used, inputs_used = dic_res["res"], dic_res["tool_used"], dic_res["inputs_used"]
if tool_used:
available_tools.pop(tool_used)
context = context + f"nTool used: {tool_used}. Output: {res}" #->add tool usage context
messages_junior.append( {"role":"user", "content":context} )
### response
agent_res = ollama.chat(model=llm, messages=messages_junior)
dic_res = use_tool(agent_res, dic_tools)
res = dic_res["res"]
print("😎 >", f"x1b[1;32m{res}x1b[0m")
messages_junior.append( {"role":"assistant", "content":res} )
El agente junior activó todas sus herramientas para explorar la base de datos y recopiló la información necesaria para generar algún código SQL. Ahora, debe informar al liderazgo.
## update Lead Agent
context = "Junior already wrote this query: "+res+ "nNow invoke the Senior to review and execute the code."
print("👩💼 >", f"x1b[1;30m{context}x1b[0m")
messages_lead.append( {"role":"user", "content":context} )
agent_res = ollama.chat(model=llm, messages=messages_lead, tools=[tool_invoke_agent])
dic_res = use_tool(agent_res, dic_tools)
res, tool_used, inputs_used = dic_res["res"], dic_res["tool_used"], dic_res["inputs_used"]
agent_invoked = res.split("-")[0].strip() if len(res.split("-")) > 1 else ''
instructions = res.split("-")[1].strip() if len(res.split("-")) > 1 else ''
El agente principal recibió el resultado del junior y le pidió al agente senior que revisara y ejecute la consulta SQL.
## Invoke Senior Agent
if agent_invoked == "senior":
print("🧓 >", f"x1b[1;34mReceived instructions: {instructions}x1b[0m")
messages_senior.append( {"role":"user", "content":instructions} )
### use the tools
available_tools = {"sql_check":tool_sql_check, "sql_exec":tool_sql_exec}
context = ''
while available_tools:
agent_res = ollama.chat(model=llm, messages=messages_senior,
tools=[v for v in available_tools.values()])
dic_res = use_tool(agent_res, dic_tools)
res, tool_used, inputs_used = dic_res["res"], dic_res["tool_used"], dic_res["inputs_used"]
if tool_used:
available_tools.pop(tool_used)
context = context + f"nTool used: {tool_used}. Output: {res}" #->add tool usage context
messages_senior.append( {"role":"user", "content":context} )
### response
print("🧓 >", f"x1b[1;34m{res}x1b[0m")
messages_senior.append( {"role":"assistant", "content":res} )
El agente senior ejecutó la consulta en el DB y obtuvo una respuesta. Finalmente, puede informar al líder que le dará la respuesta final al usuario.
### update Lead Agent
context = "Senior agent returned this output: "+res
print("👩💼 >", f"x1b[1;30m{context}x1b[0m")
messages_lead.append( {"role":"user", "content":context} )
Conclusión
Este artículo ha cubierto los pasos básicos de crear sistemas de múltiples agentes desde cero utilizando solo Ollama. Con estos bloques de construcción en su lugar, ya está equipado para comenzar a desarrollar su propio MAS para diferentes casos de uso.
Estén atentos para la Parte 4donde nos sumergiremos más profundamente en ejemplos más avanzados.
Código completo para este artículo: Github
¡Espero que lo hayas disfrutado! No dude en ponerse en contacto conmigo para obtener preguntas y comentarios o simplemente para compartir sus interesantes proyectos.
Todas las imágenes, a menos que se indique lo contrario, son del autor