Episode 23 · Development Room 404 Ep. 23: Kaio-ken
AI technical power boosts developer productivity.
<AI Kaio-ken and the meaning of power 1.1>
Panel details
Kkobugi struggles to code while Kim Saseum codes hard nearby. Labels show Power: 100 above Kkobugi and Power: 1000 above Saseum. In the corner, Park Mulbeom has Power: 0 and wears an AI-labeled sci-fi power scouter.
Kkobugi shouts 10x AI Kaio-ken! and reaches Power: 1000. Saseum is shocked, while Mulbeom admires the result through his AI scouter. Kkobugi’s shell grows huge, his eyes become red dots, and he smiles.
Saseum shouts 7x AI Kaio-ken! and reaches Power: 7000. Kkobugi stares in despair, and Mulbeom’s scouter explodes with a POP. Saseum’s eyes glow red and her antlers grow larger.
Mulbeom shouts 100x AI Kaio-ken!, but his Power remains 0. His Talk Power becomes 100000.
Intent Development Room 404 episode 23, Kaio-ken, satirizes how AI expands a developer’s technical power through a playful battle-power contest.
Episode 22 · Development Room 404 Ep. 22: Technology Race
AI has changed the rules of the AI technology race.
<The technology race rebuilt by AI 1.1>
Panel details
Kkobugi watches a distant desert race through binoculars. Cars marked MS, Google, AWS, and APPLE lead the race while countless companies ride horses behind them. Ordinary spectators watch beside Kkobugi.
Blessing beams marked AI shine down on the horses. The company riders cheer, the spectators’ telescopes become extremely long and gain AI markings, and the desert becomes a city.
The horses transform into winged robotic horses and fly close behind the leaders. AI-labeled cannons on the lead corporate cars aim at the robotic horses.
In space, Claude and OpenAI assemble and rearrange the race like a block-building game. OpenAI shines a blessing beam onto one rear section.
Intent Development Room 404 episode 22, Technology Race, imagines AI reshaping the race among technology companies in a playful four-panel webtoon.
Viewing Qualcomm only as a smartphone modem company hides both its strengths and its risks. Qualcomm combines QTL, which licenses cellular standard-essential patents, with QCT, which supplies modems, application processors, radio-frequency components, and chips for automotive and IoT systems. It is now extending the CPU, GPU, NPU, and connectivity technology in Snapdragon into PCs, vehicles, XR, and industrial edge devices for an era in which AI runs across devices.
This Qualcomm company review asks whether on-device AI can produce real diversification. Local models can reduce latency, exposure of private data, and connectivity costs, but they face memory, power, and thermal limits. Technology alone does not remove Apple and Samsung concentration, customer modem programs, Chinese demand, or licensing regulation. Figures use the FY2025 10-K and disclosed FY2026 quarterly materials. This is not investment advice.
<Official Qualcomm corporate icon 1.1>
QCT and QTL form one research engine
QCT semiconductors and QTL licensing support the same research base through different economics.
QCT sells Snapdragon, modem-RF systems, the automotive Digital Chassis, and IoT chipsets. QTL licenses inventions essential to CDMA, OFDMA, 3G, 4G, and 5G standards. Chip revenue responds to share and product cycles; licensing earns from broad adoption of cellular standards. Cash from both businesses funds another cycle of wireless and low-power computing research.
Business
Core product or right
Revenue driver
Primary risk
QCT handsets
Snapdragon, modem, and RF
Premium mix and units
Customer silicon and phone demand
QCT automotive
Cockpit, connectivity, ADAS
Design wins followed by production
Long cycles and safety liability
QCT IoT and PC
Industrial, networking, XR, PC
Edge AI and new form factors
x86, Arm, and ecosystem competition
QTL
Standard-essential patent licenses
Sales of cellular devices
Regulation, litigation, renewals
The combination is unusual. Patents arising from Qualcomm's standards work create license revenue, while its chips implement those inventions and demonstrate performance. Dependence on highly profitable QTL also creates sensitivity when regulators or customers challenge terms. A patent being essential to a standard and every commercial term being justified are separate propositions.
<Qualcomm technology platform expansion 1.2>
QCT generated most FY2025 revenue and handsets remain its largest component. Fast growth in automotive or IoT therefore does not prove that phone dependence has disappeared. Diversification becomes real when announced design wins enter production and grow large enough to absorb handset volatility.
Snapdragon coordinates CPU, GPU, Hexagon NPU, image processor, modem, Wi-Fi, Bluetooth, and security within one power budget. On-device AI can process speech, camera data, and personal context without sending every request to the cloud. Responses can be faster, raw sensitive data can remain local, and some functions can work with weak connectivity.
Yet a high NPU TOPS figure is not a user benefit by itself. Models must be quantized for limited memory; work must be scheduled across CPU, GPU, and NPU; battery life and surface temperature must remain acceptable. Developers encounter different runtimes on Android, Windows, and automotive operating systems. Qualcomm's defense is less a single accelerator than system optimization from modem through NPU plus SDKs that OEMs and developers can actually use.
Snapdragon X puts Windows on Arm performance and battery life at the center of the PC proposition. Application compatibility and enterprise management matter more than benchmark peaks alone. Snapdragon Digital Chassis combines connectivity, cockpit, and driver assistance, but automotive design wins can take years to enter production. Industrial IoT and robotics analyze camera and sensor data on site, favoring platforms that combine connectivity with efficient AI.
This positioning contrasts with NVIDIA BlueField-4, which controls data movement around CPUs and GPUs in data centers. NVIDIA expands high-power AI-factory infrastructure; Qualcomm targets the power- and thermal-constrained edge close to users. They compete in parts of the stack but are complementary when models trained in the cloud run inference on devices.
Cristiano Amon led Qualcomm's semiconductor and 5G work before becoming CEO in 2021. His central task is to reuse smartphone connectivity and compute technology across automotive, PCs, industrial IoT, and edge AI. Reuse can improve R&D leverage, but sales channels, certification, software, and support cycles differ in every market. Organizational execution cannot be assumed from technical similarity.
Qualcomm's FY2025 Form 10-K reports total revenue of $44.284 billion. QCT revenue was $38.717 billion and QTL revenue was $5.582 billion. QTL is smaller by revenue but its strong profit contribution is important to research capacity and overall economics. Cash-flow analysis must test whether Qualcomm can sustain major R&D while returning capital.
FY2025 metric
Result
What it says
Total revenue
$44.284 billion
Combined mobile recovery and diversification
QCT revenue
$38.717 billion
Chips and platforms dominate sales
QTL revenue
$5.582 billion
Patent licensing has disproportionate profit value
Key operating evidence
Auto, IoT, and PC revenue
Proof of reduced handset dependence
An automotive pipeline offers long-term visibility, but a backlog is not present revenue. Programs can be canceled and production volumes can fall, while automotive quality and safety standards exceed consumer-device requirements. PCs likewise need more than advanced silicon: Windows applications, drivers, and enterprise deployment experience matter. Related connectivity semiconductor companies show how much system value now depends on moving data reliably, not merely executing arithmetic.
Apple, licensing, AI hype, and the conclusion
Apple's expansion of internal modems is the most direct risk. A major customer moving in-house can reduce shipments and bargaining power. Samsung and Chinese handset vendors can also select internal or competing components. Customer and regional concentration make geopolitics, export controls, and demand weakness capable of arriving together.
QTL is the second risk. Fair, reasonable, and non-discriminatory terms for standard-essential patents, device-level royalty bases, and competition law have produced years of disputes. Past legal outcomes do not guarantee every future jurisdiction or renewal. Breadth is the third risk: mobile, PC, automotive, XR, and industrial IoT are attractive, but developer support and sales investment can become fragmented.
AI hype is the fourth. On-device AI must reduce cloud cost or improve privacy and latency enough to motivate upgrades. Infrequently used demos or difficult model updates will not turn NPU specifications into durable pricing. Rivals bring different strengths: Apple's vertical integration, MediaTek's cost structure, NVIDIA's AI ecosystem, and Intel and AMD's PC base.
Qualcomm's core is not one modem component. It is the ability to optimize connectivity, compute, sensing, and efficient AI as a platform. Diversification should be judged by whether non-handset revenue and software adoption reduce QCT volatility, not by the number of markets announced.
Qualcomm is using cellular patents and smartphone cash generation to move toward an edge AI platform. Automotive and PCs are large opportunities accompanied by long production cycles and ecosystem switching costs. The decisive indicators are handset revenue, Apple's modem transition, automotive production revenue, repeat Snapdragon X adoption, QTL renewals, and cash flow relative to R&D. Improvement across those measures would transform the narrative from a phone company's diversification into the expansion of an edge-computing platform.
Processors in phones, vehicle controllers, network equipment, and cloud servers need an invisible design standard if they are to share a software ecosystem. Arm is not a conventional semiconductor vendor that manufactures and sells finished chips at scale. It licenses instruction sets, CPU designs, and system IP, then collects royalties when partners ship chips containing that technology. The model gives Arm access to broad markets without owning fabrication plants, but it also creates tension because a customer can become a competitor.
This Arm company review goes beyond the familiar description of an energy-efficient mobile CPU company. It asks how Armv9, Neoverse, Compute Subsystems (CSS), and developer tools form one computing platform, and how data-center AI and Arm's move into production silicon may alter its neutral IP role. Financial figures follow FY2026 disclosures. This is not investment advice.
Arm earns license and other revenue plus royalty revenue. A chip designer pays for access to a CPU core, GPU, system IP, or an architectural license. Once products containing that design ship, Arm earns royalties according to volume and contract terms. Because a single design generation can remain in phones and embedded devices for years, earlier research and development can produce a long revenue tail.
Revenue engine
What customers buy
Recognition point
Main variable
License
Cores, architecture, and CSS rights
Contract and delivery milestones
Number and scope of large deals
Royalty
Shipped chips using Arm IP
After partner shipments
Volume, chip value, royalty rate
Software and support
Tools and engineering support
Over the contract
Developer adoption and platform breadth
The quality of the model comes from the interaction between these engines. Partner products enlarge the software ecosystem; a larger developer base strengthens the case for the next customer to select Arm. Arm reports more than 350 billion Arm-based chips shipped and over 22 million software developers. Its presence in 99% of smartphones is not merely share. It represents accumulated operating-system support, compilers, applications, validation tools, and switching costs.
<Arm license and royalty value cycle 1.2>
The two streams should not be mistaken for identical subscriptions. License revenue moves with the timing of large agreements, while royalties respond to phone and consumer-device demand and partner shipments. Related-party revenue connected with SoftBank and the structure of the China business also deserve separate attention. A high gross margin does not eliminate cyclicality or customer bargaining power.
Armv9, CSS, and Arm AGI CPU expand the boundary
Armv9 royalties are a useful indicator of adoption in higher-value products.
Armv9 strengthens security, vector processing, and AI workloads. When new cores enter premium products with higher royalty rates, mix can improve even without a dramatic increase in unit shipments. CSS packages CPUs, interconnect, memory, and validated system blocks so customers can shorten development. Customers concentrate differentiation elsewhere, while Arm captures a larger part of each design.
Neoverse is central in data centers. AWS Graviton and Google Axion have made power efficiency and customization competitive dimensions in general-purpose computing once dominated by x86. Even when accelerators such as NVIDIA Blackwell receive the attention, CPUs still prepare data and control networking, storage, and services. Arm is not trying simply to replace GPUs; it is expanding its share of host and data-movement layers around AI systems.
The Arm AGI CPU announced in 2026 marks a deeper shift because Arm is moving from supplying designs to offering production silicon. The company says it had no material FY2026 revenue impact. Success could add high-value data-center revenue and real system-optimization knowledge. Yet existing chip customers may believe their supplier is moving downstream as a competitor. The balance between ecosystem neutrality and greater value capture will define the next phase.
Software compatibility remains decisive. Server operators evaluate container images, observability tools, security agents, database extensions, and incident experience—not CPU benchmarks alone. Operational friction can slow a technically superior platform. Conversely, managed cloud services can hide architecture differences and let Arm adoption grow without end users noticing the transition.
What Rene Haas and FY2026 figures reveal
CEO Rene Haas has pushed Arm from a mobile-centered company toward a broader computing platform. The strategy combines wider architecture licensing, more complete CSS blocks, and higher value per chip in automotive, cloud, and AI. He must also protect the neutrality that motivates many competing partners to choose the same architecture.
Arm's FY2026 Form 20-F reports revenue of $4.92 billion, up 23% from $4.007 billion. License and other revenue increased 25% to $2.307 billion, while royalty revenue rose 21% to $2.613 billion. Research and development expense climbed to $2.776 billion from $2.071 billion, showing that the move into data centers, AI, and silicon carries a substantial cost.
FY2026 metric
Result
Interpretation
Total revenue
$4.920 billion
23% year-over-year growth
License and other
$2.307 billion
Larger contracts and design scope
Royalty
$2.613 billion
Better Armv9, cloud, and automotive mix
R&D expense
$2.776 billion
Investment in next-generation IP and silicon
Royalty growth was not merely a phone rebound. Arm says data-center royalty revenue more than doubled, with Armv9 and richer product mix also contributing. Manufacturing advances determine transistor economics, while Arm sells system architecture and software compatibility above them. The two layers are complementary, as shown by adjacent innovation in AI server connectivity and advanced infrastructure.
Risks, counterarguments, and conclusion
Customer concentration and related parties come first. Arm depends on product schedules and negotiations at major semiconductor and cloud customers, while SoftBank remains the controlling shareholder. China adds export controls, regulation, and local-entity complexity. RISC-V is another pressure point: it need not replace the full Arm ecosystem to reduce bargaining power in selected embedded or accelerator-control workloads.
Arm's own silicon creates channel conflict. The more successful Arm AGI CPU becomes, the more existing customers may worry that a critical supplier is learning finished-product economics. Execution and valuation are also risks. Growth and gross margin are strong, but rising R&D, stock compensation, and uneven license timing must be considered alongside non-GAAP results.
Arm's moat is not one CPU core. It is the reinforcing cycle among licenses, shipment royalties, developer tools, and operating-system compatibility. The central question is whether Arm can preserve its neutral standard while increasing value per chip through CSS and production silicon.
Arm is evolving from a “chipless semiconductor company” into a provider of a computing platform plus selected silicon. Smartphones remain the ecosystem and cash foundation; cloud, automotive, and AI provide the direction of growth. Power efficiency and software continuity are technical advantages, while partner trust is the scarcest commercial asset. The most useful indicators are Armv9 penetration, data-center royalties, CSS adoption, Arm AGI CPU customer response, and cash generation relative to R&D—not shipment volume alone.
Meta Muse Image and Muse Video matter less because they are another pair of media models, and more because of where Meta wants to put them. When AI image generation and AI image editing live in a separate web app, it mostly serves people who already know how to write prompts. Meta is trying to move that capability into everyday editing flows inside Facebook, Messenger, Instagram, and WhatsApp.
Reuters reported in July 2026 that Meta integrated Muse Image into the Meta AI chatbot and highlighted complex prompt interpretation, photo inputs, and direct edits through sketches or annotations. CNBC described Muse Image as the first image-generation model from Meta Superintelligence Labs, framing it as part of Meta's effort to reduce reliance on outside generative models and build native tools for creators and advertisers.
AI video generation has already advanced quickly in specialist production tools such as Runway. Muse is interesting for a different reason: it tests what product rules are needed when similar technology is attached not to a studio workflow, but to social feeds, messaging, and ad creation. The shift resembles the way OpenAI Codex cloud agent moved coding assistance from demos into operational developer workflows.
Background
The generative media market is splitting into two broad paths. One path is the studio-style tool: high-resolution video, character consistency, and precise scene control for production teams. The other is the everyday generative feature that uses photos, conversations, friend networks, and feed context already present in consumer apps. Muse Image belongs closer to the second path.
Meta's advantage is distribution. It does not need to wait for users to sign up for a new creative tool; it can test features directly inside large social apps. That makes the technology naturally useful for advertising drafts, lightweight memes, and repeated content variations. The downside is that consent, likeness rights, training data, and reuse of public photos become immediate product risks rather than abstract policy debates.
The Verge's July 2026 Meta archive also points to the controversy around Muse Image potentially pulling other Instagram users into AI-generated photos. Even if the feature is technically impressive, unclear rules about who can reuse which photo in what context can make trust problems surface before image quality becomes the central question.
Layer
Product meaning
What to check
Photo input
Extends an existing image into a new scene
Whether subjects and logos stay accurate
Sketch editing
Makes region-level edits easier
Whether synthetic changes are disclosed
Social distribution
Connects generation and posting
Consent and likeness settings
<Muse ad-draft use case 1.1>
Reuters and CNBC point to the same product shift: Muse is not just a demo model, but a capability placed inside Meta's own app surfaces.
How it works
As a product flow, Muse Image has four steps. First, the user enters a text prompt or provides an existing photo. Second, the model separates the subject, scene, style, and edit request. Third, it generates a new image or redraws only part of the image. Fourth, the user continues editing through sketches or annotations, such as “change only this background” or “remove this object.”
A small business, for example, could upload a product photo and ask for “the same item on a summer outdoor table, but keep the logo unchanged.” In professional editing software, that would require background removal, compositing, and color adjustment as separate tasks. In a social generative AI tool, those steps are compressed into conversational edits.
Muse Video should be read more carefully. Reuters said Meta also announced an early preview of its video generation model. Video is harder than still images because the subject has to remain stable across frames, and motion, lighting, and camera changes must make sense over time. That is why practical uses are likely to start with short clips, ad drafts, and feed-ready variations before expanding into longer production workflows.
Product structure
<Muse generative media product flow 3.1>
The left side of the diagram is user input. Text, photos, sketches, and annotations can be combined so the model can generate a new image or change a selected region. The middle safety layer represents controls for public-photo reuse, faces, brands, sensitive scenes, watermarking, and provenance. The right side is Meta's existing distribution surface: posting, ad mockups, and message sharing.
The important point is that the model is not the whole product. The same generative capability becomes a creative tool when it lives in a standalone app, but becomes part of a relationship network and privacy policy when it lives inside a social platform. Users gain convenience, while the platform has to design much stricter rules for rights, consent, and misuse.
Checkpoints
Consent and likeness: A public photo is not the same as permission for synthetic reuse. Features that call up people's images need clear defaults, opt-outs, and notifications.
Truthfulness in advertising: Fast ad mockups are useful, but they can also create scenes that do not match the real product. Disclosure and review rules have to catch up.
Temporal consistency in video: A short clip may look plausible while hands, text, logos, or object positions shift over time. Brand use still needs human review.
Platform lock-in: When generation, editing, and distribution all happen inside one company's apps, the workflow is convenient, but portability, rights metadata, and compatibility with outside editing tools may suffer.
Muse is therefore more than another image model. It shows what happens when generative AI moves from standalone services into the default editing layer of everyday apps. In that sense, social generative AI is less a separate tool category than a new interface layer for media feeds. The evaluation cannot stop at sharpness or prompt adherence; it also has to include who supplied the input, whose photos can be reused, and where the output will appear.
For readers, two practical checks matter. First, look for the settings that control whether your own photos can be used in other people's generated content. Second, if you plan to use AI-generated images for work or advertising, review rights and factual accuracy before celebrating the speed of creation.
Episode 21 · Development Room 404 Ep. 21: Department Store
The turtle searched all night.
<Google Department Store and the AI coordinator 21.1>
Panel details
On a bright morning, a shabby turtle stands outside Google Department Store and imagines himself in a neat suit.
The turtle swims frantically through a sea of mixed clothes, shoes, hats, and socks while clerks shout Shoes!, Coat!, Hat!, and Socks! beside broken items. He keeps imagining the gentleman outfit.
Late at night, the turtle leaves the store exhausted, wearing only underwear and a hat, with discarded clothes and shoes around him.
The next morning, an agent labeled AI COORDINATOR politely welcomes the same shabby turtle outside the store.
Intent Development Room 404 episode 21, Department Store, follows a shabby turtle searching through Google Department Store before meeting an AI coordinator.
Episode 2 · Learning in the AI Era 2 - The People Who Eat the Golden Apple
Part 1 described studying as recognizing a need, acquiring knowledge, embodying it, and applying it. AI does not remove this sequence; it compresses the time spent acquiring and exploring knowledge. The question in the AI era is therefore not “Did studying end because AI produced an answer?” but “Am I turning that answer into my own knowledge and capability?”
AI-assisted learning here does not mean copying an answer. It is an AI-era learning method and an AI learning method for developer growth: use the model to accelerate exploration, then attach verification and a small implementation so knowledge application actually happens.
1. AI separates learning depth, not just learning speed
AI makes it faster to locate documentation, unpack an unfamiliar term, and create a skeleton example. In the past, search felt like walking through every floor of a department store; now a coordinator can select plausible items in front of you. That is a real productivity gain, but if you do not know why the items were selected, you cannot dress yourself next season.
Developers can move in two directions:
Direction
How AI is used
What remains
Source investigator
Keeps asking for evidence, structure, and failure conditions, then verifies
Embodied design and transfer skill
Result consumer
Prompts, copies the output, and moves to the next problem
Tool-dependent manipulation
Both people may use the same AI. The difference is whether experiments and explanations follow the answer, or whether learning ends on the response screen.
2. From a Google department store to an AI coordinator
Imagine wanting to dress like a stylish British gentleman without knowing the vocabulary. A search engine can open hundreds of pages, but the user still has to discover words such as top hat, gold-rim glasses, cufflinks, wing collar, and bow tie. If you do not know the categories, you cannot even formulate the right search query.
Software problems work the same way. Search “login is broken” and you get sessions, JWT, cookies, CORS, CSRF, OAuth, proxies, and expiration policies in one pile. Before AI, developers had to walk through that department store, extract terms, combine examples that only looked similar, and discover what they did not know through compiler errors.
An AI coordinator can ask:
TEXT
Developer: Login is broken.AI: Is it a browser-cookie session or a Bearer token? Does it fail immediately or after a refresh? Is it CORS, a 401 response, or a server exception?
That exchange does not merely provide an answer. It creates categories for the problem. Once categories exist, official documentation and code searches become more precise. AI’s first value is often discovering the missing words, not producing the final sentence.
<The difference between a Google-department-store search and an AI coordinator 2.1>
3. How to question AI until you reach 97 percent
This is not a request to trust the answer. It is a request to interrogate why it was produced. When you receive code, use a question set like this:
What problem does this implementation solve, and what does it not solve?
Where are the official documents for each class, function, or protocol?
Show normal, boundary, and failure inputs separately.
Compare two alternatives and the criteria for choosing between them.
Break the idea into a 30-minute experiment I can reproduce myself.
This treats AI as a source of evidence and counterexamples, not as a teacher to obey. “Following AI to 97 percent” does not mean memorizing 97 percent of its text. It means explaining the structure in your own words, checking it with a small implementation, and correcting what fails.
TEXT
AI answer ↓ Check original sources and assumptions ↓ Write a minimal reproduction ↓ Add failure inputs ↓ Explain the design in your own words ↓ Adapt it to another problem
4. Why result consumers stop quickly
Imagine someone who asks AI for a casual outfit after receiving one satisfactory recommendation, then goes home without asking why. They have an outfit for today, but not the ability to assemble one for a different situation. Copying code, running it, and stopping is the same state.
It is fast in the short term. But when a version changes or a requirement shifts, the developer must ask the same question again and cannot tell where to begin verification. A source investigator connects a new tool to existing design principles; a result consumer starts over whenever the tool changes.
Microsoft Research’s 2025 survey of knowledge workers examined the relationship between generative-AI use, perceived cognitive effort, and critical review. It did not prove that AI necessarily reduces human thinking; it was a self-reported survey with clear limitations. It does, however, justify deliberately designing verification instead of assuming that a fluent answer has been learned. For evidence on learning strategies, compare Dunlosky and colleagues’ review with retrieval-practice research.
5. The AI learning loop of a source investigator
A deep AI user does not ask one question and stop. They decompose the problem and use AI in turn as a documentation finder, opposing reviewer, experiment designer, and code reviewer. The final judgment and execution remain their responsibility.
Stage
Developer action
What AI can do
What the developer must do
Need
Write the problem in one sentence
Find missing conditions
Set goal and scope
Acquire
Read sources and principles
Map terms, compare, summarize
Check originals and record provenance
Embody
Build a minimum project
Generate examples and failure cases
Run, log, and test
Apply
Transfer it to another platform
Propose alternatives and counterexamples
Judge quality, security, and operations
This loop turns AI from autocomplete into a learning environment. The important variable is not the number of prompts but whether each prompt connects to code, sources, and an experiment.
<Turning an AI answer into embodiment and application 2.2>
6. You have to eat the golden apple
AI lowers the entrance to knowledge that developers once found difficult to reach: operating-system APIs, unfamiliar languages, hardware documentation, and design patterns. But lowering the entrance does not make the ability yours automatically.
We are holding a golden apple. One choice is to eat it, analyze its taste and ingredients, and learn to choose the next apple yourself. The other is to leave it in the living room and keep admiring it. The first person investigates and embodies knowledge with AI; the second consumes AI-generated results.
The golden apple also decays. Today’s example becomes stale when a library changes, and an AI answer can be wrong when its context disappears. The AI-era learner therefore does not merely save answers; they turn answers into designs and execution records through verification.
Conclusion: AI changes the standard of studying, not the need to study
The sequence of acquiring, embodying, and applying knowledge remains. What changes is the speed and breadth of exploration. The advantage now lies not in receiving a brilliant answer but in connecting it quickly to sources, experiments, and design.
The future developer is not somewhere between a person who memorizes everything and a person who only prompts. It is the direction of a source investigator: recognize the need, use AI to reach the source, embody the result by hand, and apply it to a different problem.
Episode 1 · Learning in the AI Era 1 - How Knowledge Becomes Yours
For a developer, studying is not about collecting certificates. It is about expanding the range of unfamiliar problems that can be solved with your own hands. This two-part series follows the path from recognizing a need to acquiring knowledge, embodying it, and applying it to a new problem. Part 1 describes the learning structure that mattered before AI and the route by which repeated experience becomes full-stack transfer skill.
The developer learning method described here is not a trick for consuming more material. It is a process for selecting what a problem requires and turning it into embodied knowledge.
1. Learning starts with recognizing a need
The first step of studying is not opening a book. It is the moment you recognize, “I need to understand this to solve my problem.” Without that need, saving a course list or buying a book rarely becomes actual learning.
Developers begin studying when a problem arrives. A deployment fails, a data model becomes tangled, or the contract between a screen and an API breaks. The good question is not “How do I study the whole web?” but “What exact concept, API, or tool would remove this bottleneck?”
TEXT
Problem ↓Why is it blocked? → Name the required concept, API, or tool ↓Limit the learning scope to one sentence
Recognizing a need determines both direction and depth. “Study web development” is too broad; “this week, implement token issuance, validation, and expiration” is executable. The same sentence helps when using AI: it turns a stream of answers into a bounded learning unit.
2. Knowledge that flows in can flow out
The second step is acquiring knowledge. Books, official documentation, courses, videos, and other developers’ writing can all be useful inputs. But time spent watching a screen is not the same as knowledge retained in your head.
Developers especially pay a price when they skip reading for themselves. A search summary or video explanation is a starting point, not an API contract. Read the original input and output conditions, exceptions, versions, and permissions so you can explain why the code works.
After acquisition comes embodiment. To embody knowledge means you can rebuild the core from the beginning and narrow down why it fails. For developers, the best embodiment tool is often not a huge side project but a small project that can actually finish. A deliberately scoped mini project is more useful than another passive tutorial because every hidden assumption becomes an error you can observe.
If you are learning authentication, a sufficient exercise might be:
Store users in a local database.
Issue a short-lived access token at login.
Validate its signature, expiration, and permissions on a protected API.
Test an invalid token, an expired token, and a token with insufficient permission.
The screen can remain simple. What matters is starting the server, modeling data, sending requests, and reading logs and failure responses. If you only copy and run a snippet, you have observed knowledge; you have not embodied it.
<The developer learning loop from recognizing a need to applying knowledge 1.1>
4. Only embodied knowledge can be applied
Application is not pasting existing code into another project. It is redesigning a principle under new constraints. When a web developer adds a mobile app with Flutter, the work is not just memorizing new syntax; it is transferring state, events, networking, and deployment into another platform’s language.
With embodied knowledge, the question changes from “How do I draw this button?” to “Where does state live, which boundary does the event cross, and how do I recover from failure?” Without embodiment, a developer remains at the level of combining framework examples and has to search again after every small requirement change.
5. Repetition creates a specialty
Repeated projects reveal which problems you solve most often. One developer becomes strong at data models and query optimization, another at UI state and accessibility, and another at message queues and distributed systems. A specialty is not a title chosen in advance; it is the trace left by problems solved repeatedly.
Do not trade away breadth for depth. Go deep in a specialty while touching other languages and runtime environments enough to compare design choices. After several small projects, you start asking not only “Where should this feature go?” but also “Is this the right boundary for the feature?”
6. Full-stack means transfer, not a list of tools
In Korea, full-stack commonly means working on both the frontend and backend. That is useful in practice, but the heavier meaning of full-stack is closer to transferring a solution into an unfamiliar technology and reaching application level quickly.
I do not define a full-stack developer as someone who knows every technology. I define one as a practitioner who can use a focused experiment of roughly two weeks to connect an unfamiliar technology to a real task. This does not mean mastering everything in two weeks. It means finding the core contract in official documentation, building a small environment, recording failures, comparing the result with an existing design, and applying it in a narrow scope.
TEXT
New technology ├─ Find three core concepts ├─ Build the smallest runnable example ├─ Reproduce two or three failures ├─ Compare it with your specialty’s design └─ Apply it to a small real problem
This transfer skill becomes even more important in Part 2. AI dramatically speeds up knowledge acquisition, but it does not automatically perform embodiment and application for you.
<Transferring a specialty’s design across platforms 6.1>
Conclusion: the unit of study should become an artifact
Developer learning moves from recognizing a need to acquiring, embodying, and applying knowledge. Reading documentation and watching a video are necessary inputs, but a small server, screen, data model, or test must be completed before the knowledge becomes your tool. Repeated embodiment creates a specialty, and transferring that specialty across languages and platforms creates full-stack ability.
Part 2 examines how AI compresses this structure and how to use it as a partner for investigating sources rather than as an answer vending machine.
The controversy around Chinese AI models begins with a practical question: how can a smaller model answer like a larger one? Distillation is a legitimate method for model compression and transfer, but collecting a teacher model’s outputs at scale also raises API-terms and data-rights questions. Our earlier Kimi K3 distillation controversy examined that boundary, while the related AI model distillation article covers disclosed model derivatives. This article focuses on how an AI distillation technique is implemented through data and code.
The API examples below assume permission from the teacher-model provider, a valid contract, and a lawful basis for the data. They do not cover account evasion, abnormal automation, rate-limit bypass, or theft of private outputs. Use them only with an authorized API, an open model, or a model you control.
1. The foundation: what is transferred?
Knowledge distillation, formalized by Hinton, Vinyals, and Dean in 2015, transfers knowledge from a large teacher network to a smaller student network. The teacher supplies more than one correct label: it supplies a probability distribution over alternatives, a “soft target.” The student learns both the hard label and the teacher’s relationships among alternatives. The original paper made this a foundation for model compression.
LLM distillation has two broad paths. White-box distillation can access weights and logits and match intermediate representations, logits, or attention. Black-box distillation can access only an API, so it creates prompts and stores the teacher’s answers, preferences, code, or reasoning traces. The 2024 survey on LLM distillation separates these paths and discusses API-based in-context, chain-of-thought, and instruction-following distillation.
Type
Teacher signal
Student objective
Strength
Limitation
Logit distillation
Token logits and probabilities
Temperature-scaled KL divergence
Information-rich and stable
Requires teacher weights or logits
Feature distillation
Intermediate features and attention
Representation alignment and feature loss
Transfers internal representations
Difficult across different architectures
Response distillation
Text answers, preferences, code
SFT, DPO, or preference optimization
Works through an API
Copies teacher errors, style, and bias
Reasoning-trace distillation
Stepwise solutions and verification
Reasoning SFT or RL
Can improve small-model problem solving
A trace is not necessarily authentic reasoning
The basic loss can be written as follows. T is the temperature. A higher temperature smooths the teacher distribution so the student can learn which wrong answers are less implausible.
In API distillation, the student learns answers rather than logits. “More distillation data is always better” is therefore wrong. The student may also copy hallucinations, arithmetic mistakes, refusals, and verbosity. Provenance, verification, difficulty, duplication, and personal data matter before model size does.
2. Collecting distillation data: generate questions and preserve the record
2.1 A minimum schema
Distillation data collection is not simply sending prompts to a teacher and appending answers to a file. The lineage of each request and response must be stored so that the run can be reproduced, deleted, and audited later.
JSON
{ "id": "math-000001", "prompt": "A train travels 120 km in 2 hours. What is its average speed?", "teacher": "authorized-teacher-v1", "teacher_snapshot": "2026-06-01", "generation": {"temperature": 0.2, "top_p": 0.9, "max_output_tokens": 512}, "response": "Average speed is 60 km/h.", "source": "licensed-evaluation-set", "review": {"status": "verified", "method": "exact_answer"}, "hash": "sha256:..."}
First check whether the teacher’s terms permit reuse for training. If user inputs are reused, define a policy for consent, personal-data removal, and retention. Repeating the same question increases cost while reducing diversity. Sample language, length, domain, and difficulty in balance.
2.2 A rate-limited collector for an authorized API
The following assumes an OpenAI-compatible API. It reads the key only from the environment, records spacing, retries, hashes, and metadata, and must be adapted to the provider’s documented endpoint and response shape.
PYTHON
# collect_teacher.pyimport hashlib, json, os, timefrom pathlib import Pathimport requestsAPI_URL = os.environ["TEACHER_API_URL"]API_KEY = os.environ["TEACHER_API_KEY"]MODEL = os.environ.get("TEACHER_MODEL", "authorized-teacher-v1")def call_teacher(prompt: str) -> dict: payload = { "model": MODEL, "messages": [{"role": "user", "content": prompt}], "temperature": 0.2, "top_p": 0.9, "max_tokens": 512, } response = requests.post( API_URL, headers={"Authorization": f"Bearer {API_KEY}"}, json=payload, timeout=60, ) response.raise_for_status() body = response.json() answer = body["choices"][0]["message"]["content"] return {"answer": answer, "usage": body.get("usage", {})}def collect(prompts: list[str], output: Path) -> None: output.parent.mkdir(parents=True, exist_ok=True) with output.open("a", encoding="utf-8") as fp: for index, prompt in enumerate(prompts, start=1): result = call_teacher(prompt) record = { "id": f"sample-{index:06d}", "prompt": prompt, "teacher": MODEL, "response": result["answer"], "usage": result["usage"], "hash": hashlib.sha256(result["answer"].encode()).hexdigest(), } fp.write(json.dumps(record, ensure_ascii=False) + "\n") fp.flush() time.sleep(0.25) # tune only within the permitted rate limitif __name__ == "__main__": prompts = ["Explain binary search with a short Python example."] collect(prompts, Path("data/teacher.jsonl"))
This is not a bypass tool for mass scraping. In production, use the provider’s official Batch API, quotas, retention settings, and spending cap. Restrict access to raw responses because they may be sensitive, and decide whether the originals should be deleted after training.
3. Cleaning: turn responses into learnable conversations
Raw JSONL is not automatically a fine-tuning dataset. Remove empty responses, duplicates, excessively long samples, personal information, and samples without a verifiable answer. Mark only samples that pass a domain validator—equation execution, code tests, or label checks—as verified; split train, validation, and test sets by question ID.
PYTHON
# prepare_dataset.pyimport hashlib, json, refrom pathlib import Pathdef normalize(text: str) -> str: return re.sub(r"\s+", " ", text).strip()def prepare(src: Path, dst: Path) -> None: seen = set() dst.parent.mkdir(parents=True, exist_ok=True) with src.open(encoding="utf-8") as inp, dst.open("w", encoding="utf-8") as out: for line in inp: row = json.loads(line) prompt, answer = normalize(row["prompt"]), normalize(row["response"]) key = hashlib.sha256(f"{prompt}\n{answer}".encode()).hexdigest() if not prompt or not answer or key in seen or len(answer) > 8000: continue if row.get("review", {}).get("status") not in {None, "verified"}: continue seen.add(key) out.write(json.dumps({ "messages": [ {"role": "user", "content": prompt}, {"role": "assistant", "content": answer}, ], "provenance": {"teacher": row.get("teacher"), "hash": key}, }, ensure_ascii=False) + "\n")prepare(Path("data/teacher.jsonl"), Path("data/distill.cleaned.jsonl"))
The goal is not to copy every teacher mannerism. Define the target capability first. Add executable tests for a code model, a checker for mathematics, or policy and refusal criteria for support. If the prompt, verification result, and license metadata are discarded along with the teacher answer, the model’s provenance becomes impossible to explain.
<A distillation-data pipeline from question generation to verification and fine-tuning 2.1>
4. Select a base model and fine-tune the student
Choose a base model for the task and its license, not simply the largest available model. Check whether its tokenizer represents the target language and code well, whether commercial use, derivatives, and distillation are allowed, and whether context length and GPU memory fit the job. A license conflict can make a technically successful training run impossible to distribute.
Transformers provides model-specific chat templates. Do not invent role tokens; use apply_chat_template. The Transformers chat-template guide explains that models use different control tokens such as [INST] and <|user|>, and that duplicated special tokens can damage performance. The TRL SFTTrainer documentation defines preprocessing parameters such as dataset_text_field and max_length.
Start with LoRA/PEFT instead of updating every weight when GPU memory and experiment cost are limited. This model fine-tuning stage should be treated as an experiment with a held-out test set, not as proof of lineage. Check the license and merge procedure again when an adapter is merged for distribution. A falling training loss does not prove successful distillation: compare unseen questions, alternate phrasing, refusal behavior, long context, and executable code.
5. Two recent papers point to the next step
DA-KD: stop spending teacher budget on easy samples
DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language Models is an ICML 2025 paper. It argues that existing methods ignore difficulty differences and spend the same distillation cost on easy samples. Its difficulty-aware framework dynamically adjusts the dataset. The practical lesson is to spend teacher budget on what the student still cannot solve rather than simply growing the dataset.
RLKD: transfer structure beyond a flat reasoning trace
Distilling the Implicit Multi-Branch Structure in LLMs’ Reasoning via Reinforcement Learning is a 2025 study called RLKD. It argues that SFT on a single teacher reasoning trace can collapse the hidden branching of problem selection and solving into flat token imitation. It proposes a Generative Structure Reward Model to score structural alignment and uses RL to optimize it. Its results should be read within the authors’ experiments, not as a universal rule that RL distillation always beats SFT.
The papers identify different bottlenecks. DA-KD focuses on which data deserves teacher budget; RLKD focuses on the difference between copying an answer and learning a problem-solving structure. A production pipeline should design those two decisions separately.
6. How to calculate API distillation cost
API distillation cost is not the number of requests alone. Add input tokens, output tokens, cache and batch discounts, retries, teacher calls for verification, and failed requests.
The screenshot’s estimate assumes 500 input tokens and 1,000 output tokens per sample. At 200,000 samples, that is 100 million input tokens plus 200 million output tokens. Using the screenshot’s Luna example ($1 input and $6 output per million tokens), the standard API calculation is 100×$1 + 200×$6 = $1,300 (about ₩1.82 million), and Batch is $650 (about ₩0.91 million). Adding a 30–50% operating reserve gives roughly $845–$975 (about ₩1.18–1.37 million) for Batch. The $220 shown in the screenshot uses 20 million output tokens by mistake; under the stated 200,000-sample and 1,000-output-token assumption, $1,300 is the consistent result.
Actual prices vary by model, date, provider, and account. Using official prices checked on July 31, 2026, the same 200,000-sample workload is approximately as follows. The conversion uses $1 = ₩1,400 for illustration only; exchange rates, tax, and account discounts are excluded.
Model and price per 1M tokens (input/output)
Standard API
Batch (50%)
OpenAI GPT-5.6 Sol $5/$30
$6,500 (about ₩9.10M)
$3,250 (about ₩4.55M)
OpenAI GPT-5.6 Terra $2/$12
$2,600 (about ₩3.64M)
$1,300 (about ₩1.82M)
Gemini 3.5 Flash $1.50/$9
$1,950 (about ₩2.73M)
$975 (about ₩1.37M)
OpenAI lists model-specific input and output rates in its model comparison, and documents 24-hour processing with a 50% discount for Batch. Google lists separate standard and Batch rates for Gemini 3.5 Flash in its official pricing table. The screenshot’s $650 is therefore a valid low-cost-teacher scenario, not a universal price for current frontier models.
Budget for failed retries (5–10%), teacher regeneration or multiple candidates (10–30%), LLM judging and verification (10–30%), and evaluation-set generation (3–5%). For example, if Terra Batch costs $1,300 and retry plus verification overhead is assumed to be 40%, the working budget is about $1,820 (about ₩2.55 million). Use a cheaper verifier, filter easy samples before teacher calls, and apply caching and Batch where latency permits. These are approximate budgets derived from public rates and assumed token counts, not guaranteed invoices: actual charges depend on input/output/reasoning tokens, cache hits, failed requests, exchange rates, and taxes.
<A distillation experiment loop that measures teacher cost and student evaluation together 6.1>
7. Why frontier companies cannot identify every distillation attempt
Large providers cannot block every attempt because the attacker and provider have asymmetric visibility.
Legitimate use and collection can look identical at the API level. Education, evaluation, and customer support can also send many sequential questions.
An attacker can distribute work across accounts, regions, and applications and mix collection prompts with ordinary queries. The provider must avoid blocking legitimate users through false positives.
Outputs do not carry an owner’s definitive secret key. Style, answers, and refusal patterns are clues; public data and similar prompts can produce them too.
Providers may not retain every prompt and answer for long periods because of privacy, security, and contracts. The logs available for analysis are limited.
Several teachers may contribute to one student. Public data, synthetic data, human editing, and multiple models can be mixed, making attribution difficult.
Detection therefore combines abnormal query distributions, account graphs, timing, error rates, output reuse, and behavioral similarity rather than relying on one sentence watermark. Those signals can support terms enforcement and security actions, but they do not automatically prove distillation in court. Providers need rate limits, account verification, output provenance, and contractual training restrictions together; researchers need to record permission and data lineage.
Conclusion: distillation is a data contract before it is a code snippet
The heart of an AI distillation technique is not one line in SFTTrainer. It is a reproducible lineage showing which teacher received which questions, who verified the answers, and what the student was trained to learn. Distillation with permitted open-model outputs, self-generated data, and clearly licensed evaluation sets can reduce cost, latency, and deployment size. The same algorithm becomes industrially corrosive when private-service outputs are collected by evasion and their provenance is hidden.
The next competition will likely ask not who scraped the most, but who built the most generalizing student with the fewest teacher calls and the most auditable data. Distillation is neither a magic button for stealing intelligence nor a simple compression switch. It is a systems-engineering problem involving permission, cost, data quality, evaluation, and audit.