1. One project, two tellings

The most interesting coincidence of the event: 91APP sent three speakers across the two days, and two of them were describing the same project — handing the spec fields on a product listing to an AI to fill in automatically.

李昆謀 told it at the main conference from a chief product officer's vantage point, landing on how humans, a rules engine and an AI agent divide the work; 吳剛志 told it at the developer conference from a chief architect's vantage point, landing on how you take a team through dozens of improvement cycles to push the numbers up.

The two baselines do not match:

Stage李昆謀's account吳剛志's account
First version45% accuracy, 98% coverage52.3% accuracy, 87% coverage
After adding rules90% accuracy, 50% coverageOver 90% accuracy, 50% coverage
Split of the workRules engine 40, AI agent 50, human 10Roughly half AI, half human

The two decks present differently: 李昆謀's numbers are all multiples of five, drawn inside a single bar chart totalling one hundred; 吳剛志's carry a decimal place and read like raw values off an evaluation run. One possibility is that these are two presentation layers over the same data — round numbers to tell a story to a non-technical main-conference audience, precise values for an engineering audience at the developer day (inferred here, no external source). But that is an inference and cannot stand as a conclusion.

The two speakers' end states agree; their starting numbers do not. It could be different sampling batches, or two different evaluation rounds. Public sources cannot settle which pair is the project's official figure, so both are listed and neither is adjudicated. What is worth remembering is the principle both of them agree on: coverage is negotiable, accuracy is not. Better to have the AI answer half the questions than answer all of them and get seven in ten right.

2. Where enterprise adoption gets stuck

Four enterprise cases (E.SUN Bank, 91APP, Simpleinfo, STR Network) plus the logistics and printing cases describe difficulties that overlap heavily — and almost none of them are technical.

Failing the first time is normal. 林鉦育 laid E.SUN Bank's first failed rollout of an AI coding tool out in the open, concluding that "the tool got bought but nothing changed" is the default outcome, and that the deciding variables are measurement and process redesign. 黃仕鎮 likewise opened with the ugly first attempt: in the early days of the bank's internal chat platform, only 17.8% of the people it was opened to had ever used it, only 2.1% used it on any given day, and 58.8% of those who tried it never came back.

If management does not move, nobody below moves. 林鉦育 put this most bluntly; 張志祺 arrives at it from the other side: he found the real blocker was psychological rather than technical, so his team set the technology aside and came at it through experience design instead, hiding the AI inside interfaces people were already using (a Slack emoji reaction, a Google Sheets add-on) so that nobody had to "decide to start using AI".

With no off-the-shelf system, grow your own language. 鄭晴元 was working with a forty-person entertainment company for which no management system exists on the market at all. The approach was to inventory the information nodes between people and machines, and between machines and machines, then use domain-driven design and Mermaid diagrams to turn spoken-word experience into a structure that can be argued over continuously and also handed to an AI to read.

Compliance is not an excuse, it is a design constraint. Both E.SUN Bank sessions kept returning to the same sentence: customer data cannot leave, so every purpose-built application is built in-house rather than bought as an off-the-shelf integrated service. That is the exact inverse of the startup default of never rebuilding what SaaS already does.

3. Accuracy and coverage

If one pair of terms had to be picked as the most important of this edition, it would be these two. They were raised independently in three separate sessions:

  • 吳剛志 defined them as operational metrics: coverage is the share of questions the AI answers at all, accuracy is the share it gets right among those it answers.
  • 李昆謀 turned them into the basis for dividing organisational work: the AI does only the part it is confident about, and the rest goes back to the rules engine and to people.
  • 林彥廷 came at the same thing from the research end: how you design the benchmark determines which mountain everyone decides to climb.

Three people standing in completely different places, describing the same problem: how do you define "good enough" on a system that is non-deterministic by nature. 吳剛志's footnote is the most practical — a prompt you use internally can be checked case by case across a hundred runs; once it ships and runs a million times, case-by-case checking is impossible, and that difference in scale is itself the reason you need an evaluation mechanism first.

4. What agents actually get used for

Plenty of sessions covered agents, but the ones that produced concrete uses were almost all deeply unglamorous chores.

SpeakerWhat the agent actually does
Wisely ChenDownloads files on a schedule, merges them, runs a plausibility check — replacing manual reconciliation. One to two hours to build; reconciliation went from daily to hourly
保哥Takes a work item in from a form, opens a branch, writes the code and opens a PR; the human only clarifies the requirement and reviews at the end
Vivi ChenInternal posts, ad creative checks, survey processing — each squeezed from half an hour or more down to a minute
Peggy LoBookkeeping checks, form filling, administrative support — split across thirty-odd little helpers
李慕約Deep research, turning screenshots into tables (a 50% success rate by his own account, which is why he stresses learning to inspect the output)

李慕約's line — "AI agents will end up better at using tools than people are; people need to be good clients, and a good client knows how to write a brief and how to sign off" — works as the summary of this whole set. 保哥's approach is a textbook brief: the intake form deliberately displays a high per-run cost, forcing whoever files the request to state it properly.

5. Where MCP sat this year

張文鈿's session was the most technically dense of the two days, and the only one to walk a specification end to end. He was candid that remote deployment and authentication were both immature at the time, down to which version of a particular debugging tool has the bug and which version you have to downgrade to.

What is interesting is how fast the topic spread across the two days: 保哥 opened the second half by saying "ever since I saw ihower's MCP talk I changed my topic"; 黃仕鎮, in the enterprise case the following day, listed MCP in E.SUN Bank's next round of agent integration planning. A protocol that had only warmed up in the community that May turned up in the same event as a spec walkthrough, an engineering implementation, and a bank's medium-term plan.

6. How people who do not code get started

The conference deliberately programmed a whole strand for non-engineers, and all three speakers took almost the same path: build one small thing that works, then snowball.

  • Vivi Chen started with a LINE Q&A bot built in two hours; a year later she was producing a project management system with a dashboard.
  • Peggy Lo used the early-morning hours before her daughter woke up, shipping a helper every week or two over about six months. Her observation is the sharpest of the three: AI coding is a lever for the people lowest in the organisation with the most drudgery, precisely because managers have no use case of their own and therefore never learn it.
  • 海馬 compressed a six-month learning curve into one month, using nothing but four free-tier models.

All three mention the same thing: what blocked them was never "not knowing how to code", it was not knowing how to ask, and not knowing how to judge whether what the AI gave back was right. Peggy Lo says she now spends ninety per cent of her time on discussion and writing documents and ten per cent debugging — the exact inverse of where she started.

7. Two faces of AI imagery

The conference deliberately put AI imagery in the last three slots, and the order is itself a decision: first how to deceive, then how to create.

李怡志's title was "how to generate effective image deception", and the whole session was in fact about spotting it. His core claim: whether the image looks real barely matters; what matters is whether it lands on your existing position, your emotions, your knowledge gaps, and your trust in whoever forwarded it.

陳志信 and Davis Chang then discussed creation. 陳志信 argues that what decides whether a video is any good is the compositional design done before generating a single frame, not the model; Davis argues that what generative AI changed is not picture quality but who is able to start creating at all — production teams compressing from hundreds of people to a handful, after which the scarce thing is aesthetic judgement rather than tool operation.

The same day, the same hall: one demonstration of how to make a fake image land, and one of how to lower the barrier to creation, using the same set of models. That juxtaposition is the strongest piece of programming in this edition.

8. Where speakers disagree

These positions genuinely conflict, and they are not written up here as a consensus.

Should models converge or be distilled. 紀懷新's whole session is built on convergence: fold functions that used to run separately into one large model, and even for robot control he stresses one model rather than several. 吳柏翰's whole session runs the other way: distil the knowledge of a giant general-purpose model into a small model an enterprise can afford, so it can run on their own hardware. Both describe industry reality; one stands on the hyperscale platform side, the other on the enterprise-customer side.

Wait for models to get better, or patch it now with engineering. 李慕約's third strategy is literally called "lie flat" — if something is a bit of a stretch today, rather than spending effort teaching the model, wait another quarter. 吳剛志 and 李昆謀 do the exact opposite: rules engines, structured output and dozens of improvement cycles, forcing accuracy from fifty per cent up to ninety. Both are right; the difference is whether you have a product that has to ship now.

Top-down or bottom-up AI adoption. 林鉦育 and 張志祺 both locate the key in management and organisational design; Peggy Lo's observation is the reverse — the people with the most motivation and the best chance of learning it are at the bottom.

Should you look at public leaderboards. 林彥廷 spent a good while on how public benchmarks are losing their meaning, including models trading accuracy for agreeableness; 吳剛志 simply abandoned public metrics and defined accuracy and coverage for himself. The first is a researcher's problem statement, the second a product team's answer.

Should you still learn the craft. Davis Chang closed with the blunt version: even when what the AI writes is wrong, you still need an engineer, because only they have the experience to tell you whether a thing is any good — so do not stop learning. That sits in real tension with the optimistic "you don't need to code" register of the earlier sessions the same day.

9. Every quantified result, in one table

All the outcome numbers from the two days, collected here. Most are speakers' own accounts; the ones with external corroboration are marked.

SpeakerMetricFigureCorroboration
林鉦育Duplicate code densityDown over 50%Speaker's own account, no external source
林鉦育Estimated hours of technical debtDown over 40%Speaker's own account, no external source
林鉦育Share reporting clearly improved development productivity95%Speaker's own account, no external source
黃仕鎮Share of eligible staff who had used the platform early on17.8%Speaker's own account, no external source
黃仕鎮Monthly user growth after the redesign200%Speaker's own account, no external source
黃仕鎮Usage volume growth after the redesignSix-foldSpeaker's own account, no external source
黃仕鎮Recognition rate on fixed-format documentsOver 99.9%Speaker's own account, no external source
陳俊毅End-to-end delivery efficiencyUp at least 30%Speaker's own account, no external source
陳俊毅Number of system incidentsFlat year on yearSpeaker's own account, no external source
吳剛志Accuracy52.3% to over 90%Speaker's own account, no external source
吳剛志Engineer time spent validating a scoring runFour hours down to ten minutesSpeaker's own account, no external source
李昆謀Time to fill the fields by handAbout five working daysSpeaker's own calculation, no external source
保哥Share of PRs correct enough to merge as-isOver 50%Speaker's own account, no external source
保哥Cost of running a single work itemAbout US$1.26Speaker's own account, no external source
吳柏翰Gains on three benchmarks after distillation172%, 22%, 654%Speaker's own account, no external source
張志祺Rough-cut editing timeAn hour and a half down to about fifteen minutesSpeaker's own account, no external source
Wisely ChenReconciliation frequencyDaily up to hourlySpeaker's own account, no external source
Vivi ChenTime on routine tasksHalf an hour or more down to a minuteSpeaker's own account, no external source
劉詩雁Time to confirm a customer complaint responseTwo to three weeks down to an instant lookupSpeaker's own account, no external source
劉詩雁Share of the adoption workload spent preparing dataAbout 95%Speaker's own account, no external source
海馬Time to build a complete websiteSix months down to oneSpeaker's own account, no external source
紀懷新Optimisation projects handled and annual revenue contributionOver a thousand projects, about ten billion US dollars a yearConsistent with his public CV
陳志信Average attention a single post can hold1.7 secondsCited by the speaker, source not stated

The most honest thing about this table is its right-hand column: almost all of it is the speakers' own account. Outcome numbers from enterprise cases are inherently hard to audit externally. That is not a criticism of these speakers, but you should know what you are looking at.

10. Primary sources

ItemSource
Official wording of all 24 talks and speakersAgenda
The full content of each session and its own source tableDeveloper conference and the main-conference session notes
Term definitions and links to official documentationGlossary
Speaker backgrounds and press coverageSources & evidence
Speaker interviewsINSIDE 2025 Generative AI Conference feature