Góc nhìn Anderson
AI Mới Dự Đoán Phản Ứng Tiếp Theo Của Bạn Dựa Trên Các Cuộc Trò Chuyện Trước Đó

Các nhà nghiên cứu tại MIT đã phát triển một hệ thống dựa trên mô hình ngôn ngữ lớn (LLM) học các mô hình lặp lại từ vài tuần các cuộc trò chuyện thực tế của một người, để dự đoán hành vi hội thoại tiếp theo có khả năng xảy ra của họ.
Nếu AI có thể thực sự truy cập vào cuộc sống của bạn, và có thể thực sự được phép hiểu bạn trong một khoảng thời gian kéo dài hoặc thậm chí liên tục, thì nó có khả năng học cách dự đoán hành động, quyết định – và ngay cả câu nói tiếp theo mà bạn có khả năng nói; tất cả dựa trên những gì nó đã tổng hợp từ hàng ngàn giờ phân tích các tương tác của bạn.
Những người trong chúng ta đã quen dùng các Mô hình Ngôn ngữ Lớn (LLM) như ChatGPT và Claude chắc đã nhận ra xu hướng “dự đoán nhu cầu” này, rõ ràng dựa trên lịch sử sử dụng của chúng ta; và, có lẽ, đã hoan nghênh nó, vì việc ghi nhớ và ngữ cảnh hiện đã trở thành một điểm ma sát quan trọng trong việc xây dựng một mối quan hệ làm việc hiệu quả với hệ thống AI.
Tuy nhiên, việc cung cấp thông tin tạm thời và ngay cả ‘trực tiếp truyền phát’ thông tin về cuộc sống và các cuộc liên lạc của bạn cho một hệ thống AI do công ty vì lợi nhuận vận hành, có thể nói, là một rủi ro tiềm ẩn về quyền riêng tư và bảo mật; trong khi dịch vụ sẽ trở nên hữu ích hơn rất nhiều đối với bạn, sự cám dỗ cho công ty chia sẻ dữ liệu của bạn, kết hợp với nguy cơ một rò rỉ trực tuyến hoặc trộm dữ liệu, có thể khiến mức độ tích hợp AI/người này trở thành một triển vọng không khôn ngoan.
Nếu Chúng Ta Có Thể Có Những Điều Tốt Đẹp…
Thật ra, tôi phải đưa ra “cảnh báo dịch vụ công” trước khi đề cập tới một bài báo mới thú vị từ MIT, dường như tồn tại trong một thế giới tốt hơn, nơi dữ liệu kiểu này có thể được bao gói trong các vòng tròn an toàn và thực sự đáng tin cậy, để sức mạnh thực sự của AI phân tích có thể được dùng vì lợi ích thực sự của cá nhân.
Bài nghiên cứu mới đề xuất một hệ thống, với sự cho phép của người dùng và với quyền kiểm soát dữ liệu thu thập được, một chiếc đồng hồ thông minh mà người dùng đeo liên tục ghi lại các cuộc trò chuyện nhằm khai thác các mô hình hành vi:

Từ bài báo mới của MIT, minh họa quy trình dự đoán hành vi được đề xuất. Các cuộc trò chuyện được thu thập trong vài tuần cho thấy một mô hình hành vi lặp lại, sau đó được dùng để dự đoán hành vi tiếp theo có khả năng của người dùng và gửi lời nhắc trên đồng hồ thông minh trước khi hành động dự kiến xảy ra. Nguồn
Bài báo – có tựa đề Before You Say It: Anticipating Verbal Behavior from Longitudinal Everyday Conversations with LLMs – ghi lại một thí nghiệm trong đó hơn 1.000 giờ các cuộc trò chuyện tự nhiên (tức là cơ hội, thực tế) từ 14 người tham gia được thu thập qua đồng hồ thông minh, sau đó được chép lại, chuyển thành lịch sử hành vi có cấu trúc, và cung cấp cho Gemini 2.5 Pro, kèm theo các lời nhắc có liên quan đến hành vi và các quy tắc hành vi được trích xuất, để dự đoán hành vi hội thoại tiếp theo có khả năng của mỗi người dùng.
(Và giờ có thể rõ ràng vì sao bài báo mới này cần một số bối cảnh về quyền riêng tư)
Nghiên cứu mới thực chất là một thí nghiệm thực địa cho (một số) công trình trước đó của các tác giả vào năm 2026 Mind Mapper* , mở rộng công trình ấy bằng cách đưa lý thuyết vào thực tiễn, và kiểm tra xem các mô hình hành vi đã khai thác có thực sự dự đoán được hành vi hội thoại tiếp theo của người dùng trong thời gian thực, chứ không chỉ mô tả các xu hướng lặp lại.
The four authors state:
‘Kết quả của chúng tôi cho thấy các mô hình hành vi cụ thể theo tình huống, được khai thác từ các cuộc trò chuyện dài hạn, có thể cải thiện đáng kể việc dự đoán xu hướng hành vi lời nói của người dùng. Một điểm mạnh then chốt của phương pháp của chúng tôi là biểu diễn chúng dưới dạng các mô hình hành vi có thể đọc được bởi con người, có điều kiện ngữ cảnh, kèm theo các điều kiện ngoại lệ rõ ràng mà người dùng có thể xem xét và tinh chỉnh.’
‘Người tham gia đánh giá cao tính minh bạch này, [nhiều người] cảm thấy hài lòng khi có thể đọc, chỉnh sửa và đặt mục tiêu cá nhân cho các mô hình đã suy ra.’
‘Người tham gia cũng suy ngẫm về chiến lược cá nhân để thay đổi hành vi, gợi ý cách các hệ thống chủ động có thể hỗ trợ bằng cách chuyển hướng chú ý, đề xuất các lựa chọn thay thế, và đưa ra các cách diễn giải mới gắn liền với mục tiêu của họ.’
‘Điều này mở ra khả năng AI dự đoán, làm việc cùng người dùng — dựa trên các mô hình mà họ nhận ra và các chiến lược mà họ đã thấy hiệu quả.’
Phương Pháp
The aforementioned Mind Mapper is the foundation for the new system, offering a workflow designed to learn recurring behavioral rules from weeks of smartwatch-recorded conversations, rather than trying to predict future speech directly, or based on generic or prior datasets. In the Mind Mapper workflow, all recorded conversations were transcribed, anonymized, and processed through a multi-stage GPT-5 pipeline**.
(Please note that because of the protections on the older paper, we cannot, as we usually would, provide images from it, since the low-res versions publicly available are illegible)

Từ bài báo mới, minh họa quy trình dự đoán hành vi. Các cuộc trò chuyện được thu thập trong vài tuần được chuyển thành các quy tắc hành vi phụ thuộc vào ngữ cảnh, sau đó được tinh chỉnh bằng bằng chứng ủng hộ và phản bác, trước khi áp dụng cho các cuộc trò chuyện mới nhằm dự đoán phản hồi hội thoại tiếp theo có khả năng nhất của người dùng.
The pipeline generates candidate ‘if-then’ behavioral rules (i.e., rules linking conversational situations to likely responses); searches for supporting and contradictory evidence; merges overlapping rules; discards weak hypotheses; and assigns confidence and probability estimates.
Các Đối Tượng
Fourteen English-speaking adults wore the smartwatch during normal daily life for seven to ten days. An on-device voice-activity detector recorded the subjects only when it perceived speech to be present, sending 2-3 minute audio clips to a remote server for processing.
Participants received $100, and were required to inform anyone nearby that conversations were being recorded, and to obtain other people’s consent before each such interaction.
Chép Lại và Kiểm Duyệt
Audio clips were processed using Deepgram’s Nova-3-meeting model, which transcribed conversations and distinguished between speakers. After this, SpeechBrain was used as the speaker verification model, matching the participant’s voice against a 20-second enrollment sample.
The spaCy Named Entity Recognition (NER) model replaced any mentioned names, as well as other personally identifying information, with anonymous placeholders. The original audio was then deleted, and only AES-256-GCM-encrypted transcripts were stored locally on the smartwatch.
After the collection period, participants were able to review each of their own conversational transcripts through a web interface, deleting anything that they did not want included in the study, and correcting speaker labels or conversational context where necessary.
On average, only 0.16% of the transcripts were removed, and 2.48% of segments were flagged as misclassified, leaving a dataset of 15,066 utterances across the 14 participants – with 57% attributed to the wearer, and 43% to other speakers. The average utterance length was 49 words.
(However, though the paper does not touch on it, one can assume that the participants remained relatively vigilant against becoming involved in very personal verbal exchanges whilst wearing the watch.)
The transcripts were further cleaned by removing very short exchanges, as well as common speech disfluencies such as ‘err’ and ‘um’; and then merging fragmented sentences into complete utterances. This reduced the final dataset to 9,901 utterances.
To evaluate the approach, Gemini 2.5 Pro was asked to predict each participant’s next conversational response from the conversation up to that point. Its performance was then compared across three inputs: conversation history alone; condensed conversation summaries; and the new Mind Mapper behavioral rules, allowing the contribution of the learned patterns to be measured directly.
GPT-5 was used as an automated judge to score the predictions, with a separate human evaluation confirming the reliability of its assessments.
Dữ Liệu và Kiểm Thử
Test results indicate that the baseline methods tested were substantially outperformed by the authors’ pattern-based approach, with accuracy improving continuously as conversational data accumulated:

Kết quả kiểm tra so sánh ba phương pháp dự đoán. Bên trái, độ chính xác dự đoán tăng dần khi dữ liệu hội thoại ngày càng sẵn sàng, nhưng chỉ khi hệ thống học được các mô hình hành vi; trong khi các phương pháp cơ sở thay đổi ít. Bên phải, cùng một phương pháp đạt điểm cao đáng kể trên các cuộc trò chuyện mà người tham gia xác định có ý định thay đổi.
The authors contend that the system was particularly effective at predicting behaviors that participants explicitly wanted to change.
A validation study, with 40 independent human raters, confirmed the automated evaluation, ranking the pattern-based method highest in 43% of tested scenarios, compared to 24% for full context; 18% for zero-shot; and 15% for narrative summaries.
The human assessments aligned closely with the model-based scoring across communicative intent, specificity, and functional complexity, which the authors contend validates using automated judges for conversational behavior.
Two follow-up studies tested how people reacted to their mined patterns, and whether these predictions could genuinely help to change habits. Participants first reviewed their rules to flag behaviors that they wanted to break, with several returning months later to discuss the strategies they had tried, and what kinds of assistant ‘nudges’ would actually help in everyday life.
Participants flagged 114 habits that they wanted to break, spanning 912 conversational moments. When tested on these exchanges, the pattern-based method easily outperformed standard approaches by roughly 40%; and the authors contend that the system works best when anticipating ingrained habits that people actively want to change.
The authors conclude:
‘Tổng thể, các phát hiện của chúng tôi cung cấp bằng chứng rằng hành vi lời nói riêng biệt của mỗi người có thể được dự đoán từ dữ liệu hội thoại dài hạn.’
‘Điều này mở ra những khả năng mới cho các hệ thống AI tiềm năng trong tương lai, có khả năng nhận thức ngữ cảnh, dự đoán, chủ động và cá nhân hoá.’
Kết Luận
Outside of highly government-regulated applications, primarily regarding state healthcare and psychiatric scenarios, and absent all overt or tacit coercion for people to participate (such as lowered state insurance premiums, if they do)…it is hard to imagine a use of this approach that would not jeopardize the privacy and personal sanctity of the participants – especially if allied/twinned to agentic entities that could correlate recorded conversational data with digital or other types of data relating to the individual.
The MIT paper paints an interesting and vaguely therapeutic context for the approach (for instance, providing context-aware reminders about diet); but such a scheme seems more likely to attract the aggressive interest and investment of advertisers, parole and prison institutions, insurers, and a considerable number of other corporate or state entities that would like to know more about what’s on our mind.
If ever there was a compelling argument for ring-fenced, completely user-owned local AI, it’s in systems such as this, which have such severe implications in the breach.
That said, the history of advertising and surveillance tech in recent years has tended towards consumer indifference – could that lack of concern really stretch to even a system as invasive and pervasive as this?
* Một ấn bản được bảo vệ chặt chẽ hơn nhiều, không có hình ảnh độ phân giải cao công khai, mặc dù nội dung văn bản có thể được đọc bằng cách vượt qua một vài rào cản.
** Không muốn làm quá mức, nhưng bài báo cũ khó tiếp cận hơn đáng kể, vì vậy xin hãy kiên trì!
Được xuất bản lần đầu vào thứ Năm, ngày 20 tháng 8 năm 2026












