Introduction
The previous article was about thoughts after GPTs launched.
In that post, I wrote that if an Agent platform wants to be good, Tools are the key.
Think carefully about this: the GPTs we create are our products. So where is the core competitiveness and moat of the product? Is it the prompt? Absolutely not. Model capability and basic tools are provided by the platform. Prompting has no real difficulty; it is hard to defend, everyone will eventually catch up, and later the prompt itself may be simplified or weakened. Personally, I think the real moat of GPTs is still data and services, meaning customized Tools.
ElliotBai: From GPTs/GLMs Monetization: Where Is the Dawn of AI Applications?
Four months have passed. ByteDance’s Coze has stood out among many Agent platforms, which also indirectly confirms the point I made then.
What Is an Agent?
Since ChatGPT became popular, the concept of Agent has become hot again too. A lot of magical self-media commentary has kept mythologizing the term.
In reality, an Agent is nothing more than a normal model-wrapper product with an extra tool-calling layer added. Let me give an example and it should become obvious immediately.
Suppose my GPT has no tools. Then it is an angel without wings, and it is even trapped in a void.
messages = [ {"role": "system", "content": "你是个非常牛逼的Agent,你爸爸叫Elliot" }, {"role": "user", "content": "儿子,帮爸爸查查牛栏山今天的天气" }]
Now I give it a tool:
messages = [ {"role": "system", "content": "你是个非常牛逼的Agent,你爸爸叫Elliot" }, {"role": "user", "content": "儿子,帮爸爸查查牛栏山今天的天气" }]//这里开始下面就是Toolstools = [ {"type": "function","function": { "name": "GooGleSearch", "description": "谷歌牌搜索引擎,探索真实世界的开始","parameters": {"type": "object","properties": {"query": {"type": "string","description": "你想知道的问题", }, },"required": ["query"], }, }, }]
Then I give it a few more tools:
The silly son becomes a smart son.
Do Not Mythologize Agents Too Much
Anyone who has read the official docs should know that in a large model request, the two biggest variables are Messages and Tools.
Messages contain the system prompt, memory, and user query.
Tools contain capability definitions in JSON Schema.
The two together form the full prompt.
So what is the essence of Agent application development?
Dynamic prompt assembly.
Through engineering, you continuously restate business requirements into new prompts.
Short-term memory: historical QA pairs in messages.
Long-term memory: summarize the text, then put it back into the system prompt.
What is RAG?
Vector similarity retrieval, then put the retrieved content into the system prompt.
Or trigger retrieval through tools.
Action: trigger the tool_calls flag, enter the request loop, use the request parameters generated by the model to make an API request, then return the result to the large model for interaction.
When there is no tool_calls flag left, the loop ends. On the page, that corresponds to one round of conversation ending.
What are Multi Agents?
Swap out the system prompt and tools, and A becomes B.
What else is there? Nothing. At the core, that is it.
Of course, that is only the most basic principle. If you want to go deep and do it well, there are definitely many traps to step through.
How Do You Make an Agent Useful?
Why has there still been no Killer App? Why have we not seen many Agent products really land?
One reason is that Agents are unreliable. Another is that Agent developers are unreliable.
Many people worship GPT-5. Personally, I think Sam is a big storyteller.
How much can the things above really change? Does a physical attack become a magic attack?
The examples above show that the upper limit of Agent capability is strongly affected by Tools capability, which is to say, old-era business capability.
For example, if Ctrip books flights, I still need access to Ctrip’s API, right? Without an API, where do I book the ticket? Build my own Ctrip?
Beyond that, the problem is making the model choose Tools more accurately and generate API args more perfectly.
To borrow from a previous keynote, the essence of applications has not changed much.
Before, the frontend called APIs through pages.
Now, the Agent calls APIs itself.
And Then?
Workflow: design some non-general business knowledge well, then let the Agent use it directly. At this slice of time, this is the closest thing to “artificial” intelligence, and it is the most cost-effective way.
After all, you may not even know a lot of specialized business know-how. Do not expect the model to know it.
Take it slow. Keep going.
Original
This article was first published on the WeChat Official Account “白苏Elliot”: https://mp.weixin.qq.com/s/F3Ch7jE09TsKYb4FsWC-Yw