说一说我的AI产品哲学吧。
我一直好奇的是两个事物的结合:
首先是语音作为界面。人与AI交互,最自然的界面是一个语音系统。用户应该如何预期自己怎么使用它,如何在少数几步之内就被引导进入有边界的意图范围内,是一个界面问题,以前GUI回答了这些问题。语音作为新interface,需要新的解法。新界面怎么理解呢?可以分层去看,为了方便,可以暂时把语音系统类比为“软件”和“硬件”两层,语音系统的软件是系统的智能水平:意图分流和回应上的精准和便捷;语音系统的硬件是AI回复的直观体验:UI、风格、音色、长度、频次、延迟等交互细节的耦合。
然后是记录作为记忆。上下文的边界,就是一个意图被处理后的效果的边界。但这个边界有其特定的阈值,只有落在这个范围内,效果才能被控制,并带来意义。怎样的入口设计能让人低摩擦地、与当下意图本身携带的目的保持一致地进入系统,什么样的需求能让用户能持续产生新的输入,这些输入随着时间和事件的变化,怎样去控制和组织,才能带来有意义的因果性和相关性,再用这些因果性和相关性去产出效果。这些都是把上下文变成记忆时,要面对的最难的问题。
我已有的实践,都在探索其中的一些部分的排列组合,它们还非常初级。但我心里始终有一个目标,是能把这两个我最关心的方面结合在一起,做一个关于“关系”的产品,一个有历史、有循环动力的交互系统。语音作为界面,让意图进入系统;记忆作为人-人-机-机之间,数据持续连接的场,让过去和现在不断重新建立关联。意图可以被处理成意义,意义逐渐形成关系,而已经形成的关系,又会参与下一次意义的生成。
我真正想做的AI产品,就是从意图到意义的,持续的关系的发生。
Let me talk about my philosophy of AI products.
What has always interested me is the combination of two things:
The first is voice as an interface. A voice system is the most natural interface for interaction between people and AI. How users should anticipate using it, and how they can be guided into a bounded range of intents within just a few steps, is an interface question. GUIs answered these questions before. Voice, as a new interface, needs new solutions. How should we understand this new interface? We can look at it in layers. For convenience, we can provisionally think of a voice system as having two layers, analogous to “software” and “hardware.” Its software is the system’s level of intelligence: accuracy and ease in routing intent and responding. Its hardware is the immediate experience of the AI’s responses: the interplay of interaction details such as UI, style, timbre, length, frequency, and latency.
Then there is recording as memory. The boundaries of context define the boundaries of the outcome after an intent has been handled. But these boundaries have their own specific thresholds. Only within this range can the outcome be controlled and carry meaning. What kind of entry point would let people enter the system with little friction, in a way that stays aligned with the purpose carried by their intent at that moment? What kinds of needs would lead users to keep generating new input? As time passes and events unfold, how should these inputs be managed and organized so that they yield meaningful causal connections and correlations, which can then be used to produce outcomes? These are some of the hardest questions we face when turning context into memory.
My work so far explores different combinations of some of these elements, and it is still at a very early stage. But I have always had a goal in mind: to bring together these two aspects I care about most and make a product about “relationships,” an interaction system with a history and a recurring dynamic that keeps it going. Voice as an interface lets intent enter the system. Memory, as a space where data keeps connecting people with people, people with machines, and machines with machines, allows the past and present to continually form new connections. Intent can be processed into meaning. Meaning gradually forms relationships, and the relationships that have already formed then participate in the next creation of meaning.
The AI product I truly want to build is the ongoing formation of relationships, from intent to meaning.