1. In one sentence

Over the past decade Google converged two major functions, the recommender system and search, from separate rule-based systems into one generative AI model; that same force of convergence now extends to robots, and Taiwan still has room to plant a flag in combining hardware with software and in personalisation research.

2. The session

His title: three Chinese renderings coexist

The host introduced him on stage as "vice president at Google DeepMind". Checking afterwards shows that Chinese-language media do not translate his title consistently: his LinkedIn gives the English original as VP of Research, INSIDE's interview renders it as 傑出科學家, and another Chinese summary as 首席科學家. All three point at the same position with different Chinese wording, and these notes use the English original from LinkedIn, VP of Research, because that is the first-hand version on his own public profile without media translation in between (link in section seven).

Why start from history

He set out his motivation immediately: to predict the future, you first have to know what happened in the past and what is happening now. In his view Taiwan put its resources into chips from the 1990s onward and missed the software wave, when software is in fact easier to get into than hardware — he started programming at eight and had a computing doctorate from the University of Minnesota at twenty-five (both ages come from the host's introduction; his education and career were checked afterwards, see section seven).

He used Google itself as the example: as a software-centred company, its two most important lines of business are the recommender system (all the various feeds) and search, and those two functions hold up an industry worth roughly 500 billion US dollars (this figure is his own; no external corroboration was found, so it is marked as the speaker's account). Inside Google he has worked on over a thousand optimisation projects, adding about 10 billion US dollars of revenue a year — figures that line up almost exactly with the ">1000 product landings, annual revenue about $10.4B since 2013" on his public CV, so this does not look like an off-the-cuff number (link in section seven).

2015 was the turning point

He treats 2015 as generative AI's real starting point, because several things happened that year at once:

  • Image recognition error rates fell from 26% to 3%, below the roughly 5% error rate of human judgement (the roughly 5% human baseline is the long-cited ImageNet comparison, and checking external material afterwards the timing and direction broadly match, though no external source was found for that 26% starting figure — see section seven).
  • Speech recognition went from "you could barely dictate an email in a noisy bar" to being able to give commands and write messages by voice.
  • Machine translation switched from statistical methods requiring enormous phrase tables to seq2seq learning; the same year also brought feeding an image to a model and generating a caption directly.

He noted pointedly that if you had known in 2015 what was happening, "your stock portfolio would look rather different, wouldn't it" — and when he asked the room who bought Nvidia in 2015, nobody dared put a hand up.

He argues that generative AI is called "generative" because underneath it predicts one token after another, and that "generate one at a time" essence had already formed back in the seq2seq period of 2015 — nobody was calling it generative AI then. When he came back to Taiwan in 2021 and 2022 he still had not heard anyone using the term; it only became common in the press in 2023.

From divide and conquer to convergence

He says the most important lesson from his decade of research at Google is that the whole paradigm shifted from "divide and conquer" (splitting different functions into different systems) toward "convergence" — folding translation, summarisation, question answering and other formerly separate functions into a single model. He gave an example from two years ago: instructing a model on a piece of code to "find the bug in the depth first search, fix it, and change the comments into Korean", after which a colleague confirmed the Korean comments were not merely correct but rather precise — proof even then of one model handling code comprehension, debugging and cross-language generation at once.

He strung the technical lineage into a timeline: seq2seq learning in 2015, then the Transformer paper (which he called "the transformer paper everyone knows"), LaMDA which he joined in 2020, then Bard, then Gemini. He noted two papers along that road that he is credited on and thinks anyone serious about this field should read: one on chain of thought, and one he described as "the fine-tuning paper for what we now call post-training" (he did not give the exact title and which paper it is could not be established — marked unverified, see section seven).

System 1 and System 2, holding up 500 billion dollars

He borrowed the "thinking, fast and slow" framing to explain Google's two main businesses: the recommender system does System 1 fast thinking (quick intuitive judgements — a few seconds of TikTok or YouTube and you decide whether to watch), while search and coding, which need deep reasoning, correspond to System 2's long thinking. He thinks that explains why the 500-billion-dollar industry mentioned earlier concentrates in these two systems — they map onto the two ways humans process information, and on a base of 500 billion dollars an optimisation of 0.1% to 1% is worth a great deal. He also had a joke at his own expense: he has spent over a decade on YouTube's recommender system, and when he asked who in the room uses YouTube nearly every hand went up — "so when you're lying awake at night watching it, that's my fault."

Project Astra: from a phone assistant to fixing a bicycle

He played a Project Astra demo clip, in which the assistant can: identify that the component producing the high frequencies in a speaker is the tweeter; improvise creative alliteration starting with C; read a piece of code as an AES-CBC encryption and decryption function; recognise the King's Cross area of London from a frame (because it is near the station and transport hub); remember that the glasses were left on the desk next to a red apple; suggest adding a cache between the server and the database to speed things up; and connect a frame to Schrödinger's cat. He said his team built it, that the film itself was shot a year earlier, and that he had meant to bring the actual device to show but forgot to put it in his bag — describing the prototype as "only about fifty-something grams".

The new version of Astra, announced three days earlier, demonstrated a task chain at a completely different level: given "help me fix my bicycle", the assistant goes online to find the Huffy mountain bike manual, scrolls to the brake section, finds the matching repair video on YouTube, searches correspondence with the bike shop to find the three-eighths-inch hex nut size needed, phones a nearby bike shop about stock, gets sidetracked by the user asking "shall we have lunch?" and still returns to the brake passage on page 24 of the manual, and finally recommends dog basket models that mount on a bicycle to the user's requirements. He joked that he dare not lose this phone, "because when I go back to the parent company I won't be going home".

Robots: the next stop for multimodality

He believes robots are the most important next application area for multimodality. He noted Taiwan has a foundation in inventing robotic arm hardware, and that combining hardware and software is a good opportunity for Taiwan, but that three bottlenecks have long been hard to get past:

  • Generality: robotic arms used to need retraining of their vision the moment the background, foreground or object changed; with Gemini's multimodal capability they now adapt to a new setting without retraining, much as a self-driving car has to cope with endlessly varying road conditions. He played a demonstration: the arm is asked to "put the pen next to the other pencils", and to "pick up the basketball and dunk it" — for the latter he stressed the system had never seen the concept of a basketball, achieving it zero-shot.
  • Interactivity: in the demonstration clip the user keeps changing the instruction — put the banana in the transparent container, now put the grapes in the transparent container, no wait, put them in the pink container, now wipe the whiteboard — a whole run of changes also handled zero-shot. He said "a real person would have sworn long before that", contrasting it with an AI that never tires or gets fed up.
  • Dexterity: which he considers the most interesting bottleneck, covering tying shoelaces, picking up eggs, joining soft and hard materials — actions needing fine force control. In the clip the robot hand folds paper without ruining it, picks up beans with a tool without crushing them, places objects into a specific storage box, and even arranges letter blocks to spell "ace" (asked to "spell something you'd find in a deck of cards"). He singled out a small hurdle where a zip's slider has to be separated before it can be fastened, which the robot figured out for itself and most of the audience had not. He stressed the whole set of hand movements is "one model, not separate models".

Research opportunities for Taiwan

He closed by listing several directions he thinks still have plenty of research room and do not require stacking up many chips: multi-step reasoning; tool use with both physical and digital tools (he stressed that "everyone's already doing digital tool use, but physical tool use has real potential going forward"); self-improvement (teach it once and it keeps learning); multimodal reasoning input and output; and personalisation and customisation. He thinks Taiwan has plenty of opportunity in all of these.

3. Figures and cases

  • Over 1,000 optimisation projects handled, adding about 10 billion US dollars of revenue a year to Google — consistent with the "annual revenue ~$10.4B since 2013, >1000 product landings" on his public CV (see section seven).
  • The recommender system plus search hold up an industry of roughly 500 billion US dollars (speaker's own account, no external source).
  • In 2015 image recognition error rates fell from 26% to 3%, below the human rate of 5% (direction and timing match external material; the 26% starting figure is unverified — see section seven).
  • He is a co-author on the chain of thought paper (checked afterwards, see section seven).
  • He joined LaMDA in 2020, which evolved into Bard and then Gemini (LaMDA's paper year checked afterwards, see section seven).
  • The Project Astra prototype weighs about fifty-something grams, with the demonstration clip shot a year earlier; the new version announced three days earlier runs on a phone and completes multi-step chains — manual lookup, finding a video, searching email, making a phone call (Project Astra's official overview checked afterwards, see section seven).
  • The three robot bottlenecks — generality, interactivity, dexterity — are his own framing (checking Google DeepMind's account of Gemini Robotics' three capabilities afterwards, they correspond almost exactly to his three bottlenecks — see section seven).

4. Lines worth keeping

  • "Why do you need to know what's happening now? So that you can predict what's gonna happen in the future."
  • "What we're talking about now is convergence — converging all the functions, all the models, into one model."
  • "But my classmate who's more famous than me, who went abroad to study the same year, is 林志玲. 林志玲 was my classmate."
  • "This is completely one model, not separate models, and it can control the whole dexterity of the hand."
  • "Taiwan invents a lot of robotic arms; right now there are three bottlenecks that have in fact been hard to get past."

5. Tools and terms mentioned

  • Gemini: Google's current multimodal generative AI model family, and the product at the centre of this talk's "Gemini Era".
  • Bard: the name of the conversational product that preceded Gemini.
  • LaMDA (Language Model for Dialogue Applications): Google's conversational language model from the early 2020s, whose technical line evolved into Bard and Gemini.
  • seq2seq learning (sequence-to-sequence learning): the neural architecture mapping one sequence directly to another, the key technology in machine translation's turn away from statistical phrase tables.
  • Transformer: the neural network architecture proposed in 2017, the foundation of contemporary large language models.
  • Chain of thought (CoT): the prompting technique of having a model emit intermediate reasoning steps to improve problem solving.
  • Post-training / fine-tuning: the stage of tuning a model after pre-training is complete.
  • System 1 / System 2: psychology's dual-process theory of fast (intuitive) and slow (deliberate, effortful) thinking, which he borrows to explain the difference between recommender system tasks and search/coding tasks.
  • Project Astra: Google DeepMind's research prototype for a universal AI assistant.
  • Gemini Robotics: Google DeepMind's model extending Gemini's multimodal capability to robot control.
  • Generality / interactivity / dexterity: the three key capabilities he identifies for robotics research to cross.
  • Zero-shot: a model's ability to complete a task without task-specific training.
  • ImageNet: the large image recognition benchmark dataset, which he uses to illustrate deep learning overtaking human judgement in 2015.

6. Wider observations

The clearest methodology in the session is his use of a sense of history as a tool — not nostalgia, but calibrating judgements about the future against what has already happened, which is more instructive than any demo. His insistence on "it had already happened in 2015" is really a reminder to the developers in the room: capabilities that look self-evident now usually leave traces three to five years earlier, and what matters is whether you noticed at the right moment.

The three robot bottlenecks (generality, interactivity, dexterity) are the most structured part of the session, turning a vague "what robots still cannot do" into three dimensions you can check one at a time — no wonder he connected them to Taiwan's hardware and software background. If those three bottlenecks really do open up one by one, Taiwan's hardware foundation in mechanisms and sensors does have a chance of meeting the model capability on the software side.

7. Sources

ItemSource
紀懷新's current role as VP of Research, Google DeepMindLinkedIn
Education: three degrees from the University of Minnesota (BA 1994, MA 1996, PhD 1999), completed within 6.5 yearsedchi.net resume
Inconsistent Chinese renderings of his title (研究副總裁 / 傑出科學家 / 首席科學家)INSIDE interview
The public CV's ">1000 product landings, annual revenue ~$10.4B since 2013"LinkedIn, Google Research profile
His COMPUTEX 2025 talk title and slotCOMPUTEX Keynote & Forum official page
Co-authorship of the chain of thought paperarXiv 2201.11903
The LaMDA paper (2022)arXiv 2201.08239
ImageNet error rates overtaking human judgement in 2015 (human baseline about 5%)aiimpacts.org
Project Astra as a universal AI assistant research prototypeGoogle DeepMind page
Gemini Robotics' three capabilities: generality, interactivity, dexterityGoogle DeepMind blog