THEME:
guide@ai-foundations:~/

$ cat AI_FOUNDATIONS.md

AI Foundations

Understanding the systems you are working with

Guide version: 2026-10-01. View this edition's frozen snapshot · Version history

You ask an AI assistant to prepare a customer follow-up. It writes a useful message. Then you ask it to check the latest order, remember a preference and repeat the work next week. Those requests sound like a natural continuation. They require several different parts of a system to work together.

This guide explains those parts: how models produce answers, what information they can use, how context and reasoning consume resources, and how connections and permissions turn an assistant into something that can do useful work. It also explains the practical choices between a hosted service, a model on your computer and a company-operated system, including hardware, speed and cost.

It is written for students, entrepreneurs, managers, employees and technical practitioners. You do not need to code, buy a subscription or read another guide first. We follow one fictional customer, Northbank, from a simple draft towards a repeatable process. Each example tells you which records are available and what you have asked the assistant to do.

By the end, you should be able to explain why an assistant behaved as it did, weigh where to run a model and make a short sketch of an AI system: its purpose, sources, tools, identity and what counts as done. That gives you something concrete to discuss with a colleague, supplier or technical team.

The foundations should outlast a particular model launch. I separate general mechanisms from provider-specific behaviour and dated examples. Exact capacities, settings and prices change; the distinctions between learned capability, available information, permitted action and verified result remain useful.

I develop this guide through teaching, practical experience and reader feedback, and update it as the explanations and evidence improve. The main guide carries the latest edition. Dated snapshots preserve what earlier readers and classes used, while the version history explains what changed.

This is a foundation with practical starting points, not an exhaustive product manual, a model ranking or a complete treatment of AI governance. The related guides develop software delivery, bot-managed work and security in greater depth.


Part 1: Learned capability

1. The model and the application

For the customer follow-up, an AI model may know how such messages usually sound. It does not thereby know what you promised this customer yesterday or have permission to send a message. To see why, separate the learned model from the application around it.

Artificial intelligence covers systems used for tasks such as prediction, classification, generation and planning. A model that estimates tomorrow's demand and one that drafts an email are both AI, but they do different work. This guide concentrates on the generative models and connected assistants people increasingly use at work.

A model is a learned computational system that transforms inputs into outputs. During training, examples and feedback are used to adjust its internal numerical values, called parameters. Many of these values are called weights. Together with the model's computational structure, they determine how it processes information. Learned weights can encode patterns and memorised training material, including text a model may reproduce. They are different from a searchable collection of current customer records.

Further training after the initial learning stage is called post-training. It can shape instruction following, style and when a model refuses a request. These training stages change learned behaviour; the AI Security Primer explains why behaviour and enforced access are separate protections.

Using a trained model to process a new input is called inference. A large language model, or LLM, generates language by producing sequences of tokens: units of text that we examine in section 2. In a common generation process, it predicts possibilities for the next token from the available input and the tokens already generated. The system selects a continuation and repeats the process. It can also produce convincing statements that are unsupported or wrong, commonly called hallucinations.

Why does that happen? Producing a plausible continuation does not, by itself, check each claim against reality. When information is absent, ambiguous or misunderstood, the model can fill the gap with a familiar pattern instead of recognising the gap. Training and evaluation can improve factual behaviour, but fluent language remains different from verified evidence. Relevant sources, suitable calculation tools and checking important claims can reduce errors; asking for confidence alone cannot establish accuracy.

For example, imagine asking When is Northbank's delivery due? before supplying any record. An answer such as It is scheduled for Tuesday would sound useful but have no basis in the supplied information. That is an illustrative failure, not a measured response: an assistant might instead correctly ask for the order record. The important distinction is whether the date has a source.

The model runs inside an application: the software that accepts your request, supplies instructions and information, calls the model and presents the result. The name on a chat window does not describe the whole system behind it.

1. Your purpose

Prepare a useful customer follow-up.

2. The application

Assembles instructions, available context and tools.

3. The model

Generates a response or a request to use a tool.

4. The system result

A draft, retrieved record or completed action to check.

Two applications using the same model can therefore behave differently. Think of asking about yesterday's news: an assistant with web search can retrieve recent reporting; one without current sources cannot establish what happened yesterday from old training alone. A visible searched the web step is the application obtaining information, not the model suddenly acquiring new training. The organisation operating a model service or application is its provider; these parts can have different providers.

Useful starting jobs include drafting a message, summarising supplied material, extracting dates into a table, translating, explaining, comparing documents and generating code. The model contributes learned patterns of language and structure. The job determines what sources, tools and checks are needed to make that contribution useful.

A prompt is an instruction or other input supplied to guide the model. A system prompt supplies overarching instructions, such as the assistant's role and operating rules, separately from an ordinary user message. Platforms use different message roles and priorities. The application may also add business context and output requirements that are not visible in the chat.

Give the instruction a useful shape

You manage customer orders and want a draft delivery update for a customer called Northbank. You have supplied the assistant with this current order-record excerpt:

  • Customer: Northbank.
  • Confirmed delivery: Tuesday.
  • Next step: ask the customer to confirm someone can receive the delivery.

A vague instruction is: Follow up with the customer. Even with the record available, it leaves the deliverable and permission to act unclear.

A clearer instruction is: Prepare a short customer follow-up using the order record we have just reviewed. State the confirmed delivery date and the next step. If either is missing or conflicting, ask me rather than inventing it. Return the draft, followed by the facts you used and any unresolved questions. Do not send anything.

You have named the audience, purpose, source and result. An example can make the desired format clearer: Use this shape: Draft: ... / Facts used: ... / Questions: ... The example shows structure, not customer facts to copy. Once the facts are right, you can ask for a warmer or shorter draft without starting the task again.

Try it: supply the fictional record above, then compare the vague instruction with the clearer one. Next ask for the three labelled parts. Check whether the result separates the customer message from its supporting facts and questions, rather than judging only which sounds more polished.

Architecture and training

A model's architecture describes how its computation is organised. Many language models use transformer-based architectures, in which attention mechanisms help relate parts of the input. Attention is a mathematical operation; it is not evidence of human awareness.

Initial training learns broad patterns. Fine-tuning adapts an existing model using further training data. One post-training technique is reinforcement learning from human feedback, or RLHF, which uses human preferences to guide a reward signal for training. Other approaches use supervised examples, AI feedback or other rewards. RLHF is not a description of every model's training process.

Distillation can train a model using outputs or other supervision from another model. These methods change learned behaviour; supplying a document in a conversation usually changes the current input, not the weights.

Further explanation: Google's language-model introduction and fine-tuning and distillation overview. Research demonstrates training-text memorisation and describes instruction-following training with human feedback.

You can now separate three things in a follow-up: the model can draft the wording, a record must supply the commitment, and a connected service would be needed to send it. Writing a convincing message establishes only the first.


Part 2: Available information

2. Tokens and the context window

A tokeniser converts text into units called tokens, represented by numbers. A token may correspond to a word, part of a word, punctuation or another unit. Token counts depend on the tokeniser and content: a page of prose, source code and a table of identifiers need not occupy the same amount of space.

For a sense of scale, the four paragraphs in the contract-upload example below contain 177 space-separated words and measured 202 tokens with one tokeniser. A short question and answer might use only a small part of a window; long documents, repeated history and tool results can use much more.

Measured October 1, 2026 with tiktoken 0.12.0, cl100k_base: plain paragraph text joined by blank lines, excluding the heading, HTML and application instructions. This is an example, not a words-to-tokens conversion rule.

Different languages, writing styles and tokenisers give different counts. Use the relevant system's counter when capacity or cost matters. Hugging Face explains tokenisation; OpenAI's counting example describes the tool used for this measurement.

The context is the information made available for the current model operation. The context window is the working capacity available; its maximum is the context-window limit, usually expressed in tokens. Think of a working set, not an archive of everything the application has ever stored. Many systems must accommodate both supplied context and generated material within this limit, with a separate output-token limit for generation.

A message such as this conversation is too long is one way an application can expose a limit. Another may manage the history without showing that message. The screen can still display older messages that are no longer included in the model's next working set.

The visible chat is not a reliable inventory of that working set. Depending on what the application supplies, space can be occupied by:

  • Application instructions, your current message and selected conversation history.
  • Earlier answers, retrieved passages, uploaded document content and saved preferences.
  • Descriptions of available tools, tool requests and returned results.
  • Images, audio or other supported inputs, represented and accounted for according to the model.
  • Reasoning and other generated material that the particular system retains.

In our follow-up, the instructions, retrieved order record, generated draft and descriptions and results of lookup tools can all occupy the window. A million files on a connected drive do not all enter it because the drive is connected. Conversely, one tool that returns a huge log can consume substantial space even when your request was one sentence.

A conversation uses more context than the chat bubbles show. You are preparing Northbank's delivery update. A tool request asks the surrounding software to look up an order; the returned record is then supplied to the model. In this illustration, the application keeps the earlier messages and lookup results for the next exchange. Each snapshot shows the material present at the end of a response, not everything the application has stored.
InstructionsUser messageTool requestReturned informationReasoningAssistant reply
1. Ask for the delivery date

Context window: the same capacity limit at every stage

Supplied to the model for the reply, after a lookup

  1. System instructions and tool descriptionsUse current order records. Prepare drafts only. A tool can look up an order.
  2. User messageWhen is Northbank's delivery due?
  3. Earlier tool requestLook up Northbank's current order.
  4. Returned order informationConfirmed delivery: Tuesday.

Generated in this model call

  1. Assistant replyThe order record says Tuesday.

Unused capacity remains.

2. Ask for a customer message

Context window: the same capacity limit at every stage

Supplied for the next reply, including retained history

  1. System instructions and tool descriptionsUse current order records. Prepare drafts only. A tool can look up an order.
  2. Earlier user messageWhen is Northbank's delivery due?
  3. Earlier tool requestLook up Northbank's current order.
  4. Earlier returned informationConfirmed delivery: Tuesday.
  5. Earlier replyThe order record says Tuesday.
  6. New user messageDraft a short update asking them to confirm someone can receive it.

Generated in this model call

  1. Assistant replyYour delivery is scheduled for Tuesday. Please confirm someone can receive it.

More material is present; the capacity limit has not grown.

3. Check again before using the draft

Context window: the same capacity limit at every stage

Supplied for the reply, after another lookup

  1. System instructions and tool descriptionsUse current order records. Prepare drafts only. A tool can look up an order.
  2. First user messageWhen is Northbank's delivery due?
  3. First tool requestLook up Northbank's current order.
  4. First returned informationConfirmed delivery: Tuesday.
  5. First replyThe order record says Tuesday.
  6. Second user messageDraft a short update asking them to confirm someone can receive it.
  7. Second replyYour delivery is scheduled for Tuesday. Please confirm someone can receive it.
  8. New user messageCheck the latest status before I send this.
  9. New tool requestLook up Northbank's order again.
  10. New returned informationDelivery delayed. Replacement date not yet confirmed.

Generated in this model call

  1. Internal reasoning, where usedWork on the response; this is not necessarily visible in the chat.
  2. Assistant replyThe latest record reports a delay. Confirm the new date before using the earlier draft.

The instructions, retained exchanges, lookup results and current generation all use capacity.

Illustrative: blocks identify kinds of content, not token counts; one exchange can involve several model calls. Reasoning is shown only in the last snapshot for clarity, while the reasoning used and history retained depend on the system. Stored files and messages enter the window only when their content is supplied.

Text brought into context can also contain attempted instructions. Treat a retrieved email as information about the task, not permission to change it. The AI Security Primer and OWASP's guidance explain this influence problem, called prompt injection.

Multimodal means a system can work with more than one kind of input or output, such as text and images. It does not mean every model can process every file type. An application may extract text from a PDF, render its pages as images, retrieve selected passages or reject it. Those routes differ in cost and in what information survives.

Why "I uploaded it" is not enough

You manage a customer's annual service renewal. You upload the contract to an assistant and ask: Does this agreement allow us to increase the price at the first renewal?

In this illustrative setup, the application stores the file, searches it and supplies selected passages to the model. It supplies the general clause allowing an annual increase but misses an appendix freezing the price for the first renewal. The assistant answers that an increase is allowed, using the incomplete material it received.

The exception was in the uploaded file, but not in the model's current context. The resulting draft could quote the customer the wrong price. Uploading made the file available to the application; it did not establish that the model considered every relevant provision. Other applications may supply the whole document or process its pages differently.

Before relying on the answer, inspect the cited passages and the relevant appendix, then have the answer checked against both. The assistant's claim that it read everything is not proof. More confident wording cannot compensate for an exception it never received.

A larger window can help with long material, but capacity is not a guarantee of effective use. Relevant facts can be obscured by repeated, conflicting or unrelated content. When an answer deteriorates, consider what the application supplied before assuming the model needs still more text.

Provider documentation: Anthropic's context-window explanation describes the contributions of instructions, tool use and thinking.


3. Memory and compaction

You finish preparing a customer update today and return tomorrow. What can the assistant carry forward? Its learned weights still provide capabilities acquired during training. The next response depends on the context the application supplies now. Information saved outside that context, such as your earlier conversation, preferences or working notes, has to be brought back into use.

Persistent memory is information kept for later use, usually by the application or an external service. A memory settings page may let you inspect or remove saved preferences. Those settings concern stored information, not the model's learned weights. To influence a later answer, the relevant information must be brought into context.

An ordinary inference call does not retrain the model with your conversation. The application creates continuity by bringing back selected history, saved information or working records. It may store these itself or use a provider's conversation service. It can also shorten what it supplies, which is why the saved chat and the next working context can differ.

Files that carry instructions and working records

A file ending in .md is a Markdown document: plain text with simple formatting for headings, lists and other structure. People and AI applications with file access can read and update it. The file can hold instructions, saved facts or working notes; its purpose depends on what you put in it and how the application uses it.

  • AGENTS.md: how to work in a project. It might say where to find current records, how to check changes and when to ask before publication. Codex and other supporting tools recognise this filename for project guidance.
  • MEMORY.md: information intended for later use. It might record that you prefer concise customer messages. Whether this filename is loaded automatically depends on the application and its configuration.
  • project-state.md: where the work stands. For our follow-up, it could say: Draft saved. Delivery date needs confirmation. Sending is not approved. This is an example filename chosen for the project, not a universal standard.

These are examples, not required filenames or a document hierarchy. Some products keep comparable information in databases or built-in memory features. A supported instruction file may load automatically; another document may need to be requested. Check the application's loading rules. Saving a file preserves its contents; only content brought back into context can guide the next response.

Provider documentation: AGENTS.md project instructions.

What compacting does

Compaction reduces the working context while carrying forward information intended to support continued work. Often this involves a summary; some systems use a specialised representation that is not ordinary readable prose. It does not enlarge the model's context window. It makes room within it.

Do not assume compaction is a lossless copy of the conversation. A qualification, rejected option or reason behind a decision may be omitted. Summaries can also preserve a mistaken interpretation. The old conversation may remain in the application's storage while no longer being available in full to the model.

For work that matters, keep the current objective, confirmed decisions, important evidence and unresolved questions in dependable records. Customer commitments, permission to send and unresolved disagreements should not exist only in a conversational summary. After compaction or a new session, retrieve what the next task needs. A fresh conversation removes available context as well as clutter; give it the current task and records rather than merely saying continue.

A preference is not a promise

Yesterday you told the assistant to keep customer messages short. You also discussed Friday as a possible delivery date, but nobody confirmed it. Today the order record contains an approved Tuesday delivery commitment.

A saved preference can supply the short writing style. The current order record must supply Tuesday. A summary saying discussed Friday delivery is not permission to promise Friday. If the current record cannot be obtained, the assistant should flag the missing confirmation instead of turning yesterday's discussion into a customer commitment.

Try it: in a chat using fictional information, give the assistant a writing preference and the Northbank record. Ask for a draft, then make the same request in a fresh chat. Note which information is available in each: a new chat may still receive saved preferences. Supply the record where it is missing and compare the result. The lesson is to inspect what carried over, not to assume that every fresh chat starts empty.

Provider documentation: OpenAI conversation state, OpenAI compaction and Anthropic context management.


4. Reasoning, budgets and cost

Your application may offer a faster mode, a thinking mode or a choice of models. What are you changing? A model choice can change learned capability. A thinking control can change how much computation a supported model spends working through a problem. A product's mode may change several settings at once, so check its explanation rather than assuming the label describes one mechanism.

Additional thinking is often exposed as reasoning effort, thinking time or a reasoning budget. It can help with comparing approaches, tracing dependencies or checking a proposed solution. For Northbank, it might help untangle conflicting delivery records. It cannot recover yesterday's promise if no record of that promise is available.

Temperature and reasoning effort do different jobs

Temperature, where available, changes how the model selects among possible continuations. Lower values favour more probable continuations more strongly; higher values allow more variation. For a customer message this can change wording, but it can also change substance or quality. Temperature is not a truth setting.

Reasoning effort concerns work spent on the problem. If two delivery records disagree, extra reasoning may help compare their status and evidence; increasing temperature does not request that comparison. Neither control supplies a missing record. Some applications expose these choices, while others manage them for you.

How temperature changes generation

The model assigns probabilities to possible next tokens. Sampling selects a token using those probabilities; temperature changes how concentrated the selection is. It does not change learned weights. Even a low setting is not a universal guarantee of identical replies: other settings and the execution system also matter.

Provider documentation: generation and sampling controls and thinking controls.

Which limit are you changing?

The word budget needs a qualifier. These limits answer different questions:

Capacity, work and spending are different limits
LimitWhat it boundsFor the follow-up
Context-window limitThe current working capacity.Room for instructions, records, retained messages and generation.
Output-token limitGenerated material for a response, including reasoning in some systems.The response can stop before the account of commitments is complete.
Reasoning budgetInternal reasoning work where a budget is supported. An effort level need not be an exact token cap.Work spent comparing conflicting records before answering.
Task budgetA whole task, potentially involving many calls and tool operations.Resources allowed for looking up, drafting, checking and retrying.
Monetary budgetSpending, as defined and enforced by the application or service.How much the work may cost, not how much the model can see at once.

An API, or application programming interface, is a defined way for one piece of software to request information or actions from another. Many model APIs count generated reasoning tokens within the output-token limit as well as the context used during processing. Increasing reasoning effort can therefore use more time, tokens and available capacity without making the final answer longer. Whether earlier reasoning is included in later requests depends on the model and application; it is not safe to assume it is always kept or always discarded.

One window, several uses

A team asks its assistant to compare a long customer correspondence with the order history and produce a supported account of outstanding commitments. The application supplies the messages, records and instructions; the model must work through disagreements before writing its account.

Calculated illustration, not a model specification: imagine a system with a 100,000-token context-window limit and a 25,000-token output-token limit that includes reasoning. It receives 70,000 tokens of instructions, history and source material. A response using 20,000 reasoning tokens and 5,000 visible-answer tokens reaches the output-token limit, even though 5,000 tokens of window capacity remain.

70,000: supplied context

20,000: reasoning

5,000: visible answer

5,000: unused window capacity

If the response is unfinished when it reaches the output-token limit, generation stops. The team might receive only part of its requested account even though the context window still has space. In a metered API, work already consumed can still be charged even if no useful answer is returned. More room in the window does not override the output-token limit, and a higher reasoning setting does not enlarge either limit.

Now imagine the application completes the task through several model calls. The total tokens processed over that task can exceed the size of any one context window. Each call has its own working set, and earlier material may be repeated, removed or compacted. A task budget of 300,000 tokens does not imply that the model can see 300,000 tokens at once.

The relationships above are illustrated by OpenAI's reasoning documentation. Check the selected model's accounting rather than transferring a setting or limit from another provider.

Why a short answer can cost more

The approved customer email may be only three sentences, but preparing it can involve retrieving several records, retrying a failed lookup and comparing conflicting information. The visible answer is only part of the work. Usage can include input processed again on later calls, reasoning, output and separate tool charges. A flat subscription may hide those components behind usage allowances rather than showing an API bill.

Latency is the delay before or during a response. Throughput describes how much work a system processes over time. Fast token generation helps, but a task may spend most of its time waiting for a database, a browser, another service or a human decision. Measure completion of the useful task, not just the speed of words appearing. The deployment track explains tokens per second, the wait before the first token and the hardware behind those measurements.

Three different things called a cache

A prompt cache reuses processing of a repeated input prefix. It can reduce cost or processing time, but the reused input still forms part of the model's context. A KV cache is a different mechanism inside the model's execution engine: its practical cost is working memory that can grow with context.

An answer cache stores a completed response for reuse. That can avoid another model call, but the response can become stale or be inappropriate for a different user's permissions. These caches solve different problems. None is a general promise that the assistant remembers everything.

Technical references: prompt caching and KV cache strategies.

When is more reasoning worth trying?

A stated fact: the supplied order record says Confirmed delivery: Tuesday, and you ask the assistant to return the date and its source. Try the default or a lower supported effort setting and check whether it extracts Tuesday correctly and identifies the right record. Then try a copy with the date left blank: it should flag the missing confirmation rather than guess. If a higher setting produces the same checked results but takes longer, the extra work has not helped these tasks.

Conflicting evidence: an earlier customer email promises Tuesday, a warehouse note says Thursday is possible, and the order record still says Tuesday. You ask what the customer can safely be told. More reasoning is worth testing here: can the assistant distinguish a confirmed commitment from a tentative operational estimate, identify the unresolved conflict and prepare the question for the person who can resolve it? The desired result is not simply to choose the newest date. Until the business confirms a change, extra thinking cannot turn Thursday into an authorised promise.

Compare settings on the same supplied material and the same task. Include a straightforward case, a conflict and a missing record. Note whether the answer is usable after checking, how long you waited, any visible usage cost and how much correction you had to do. If additional effort consistently catches an important conflict that the quicker setting misses, it may be worth the delay. If both fail because the order record was never supplied, fix the information path rather than increasing effort again.

Use the same comparison when choosing between models or a faster and a thinking mode in your app. Change one choice at a time where the interface allows it, and record the model or mode used. The useful choice is the one that handles your ordinary cases and important exceptions at an acceptable total cost, including your time correcting it.


5. Working with business information

Useful business work often depends on facts that were never in a model's training data: today's order status, the current contract or an internal decision. Retrieval obtains relevant material from a source and supplies it for the task. Retrieval-augmented generation, usually shortened to RAG, combines that material with generation.

Instead of asking someone to assemble every fact before requesting a draft, the system can retrieve the relevant account context and use it to prepare the message. It still needs to find the right information, recognise its date and scope, and preserve important qualifications.

A training cutoff describes a boundary in the training information used for a model. It is not a guarantee that the model knows every earlier fact or has a live view of later events. The application can supply today's date, but knowing the date is not the same as knowing what happened today. Web search is a form of retrieval: it can bring newer evidence into context without changing the model's weights. That evidence still needs checking for relevance, reliability and freshness.

An authoritative record is the record the organisation relies on for a defined fact or decision. A model can summarise it, calculate from it or draw an inference, but the result should distinguish recorded facts, calculated results and inferred conclusions. A generated summary does not become a competing original. Source links help readers verify claims, but the presence of a citation does not prove that the cited passage supports the answer.

An embedding represents information as a list of numbers that can support comparisons. A search system can use embeddings to find passages with related meanings, even when their wording differs. A vector database stores and searches such representations. Retrieval may also use ordinary keyword search, filters or a direct database query; RAG does not require one particular database product.

Similarity is not authority. Searching for Northbank's payment terms might return an earlier proposal offering 60 days, while the signed agreement says 30. The proposal's wording matches the question, but it does not establish the agreed terms. Retrieval needs the right customer's current, applicable source and enough surrounding text to interpret it.

For a customer follow-up, take the facts from the appropriate records and use the model to help express them. The distinction matters when several business systems describe different parts of the same transaction.

The follow-up gets its facts

You want to update Northbank about delivery and acknowledge payment if it has arrived. The customer relationship management system, or CRM, records that Northbank accepted the proposal. The order system records Tuesday as the promised delivery date. The accounting ledger shows the invoice is still unpaid.

A draft saying Thank you for your payment; delivery is Tuesday would confuse an accepted order with received money. The assistant needs the order record for delivery and the ledger for payment status; the CRM alone does not answer both questions.

If the customer reports payment but the ledger still shows unpaid, retain that discrepancy and ask the responsible person to check it. Do not turn conflicting records into a confident statement that nobody has verified.

Adding knowledge or changing behaviour

Fine-tuning further trains an existing model on additional examples. It addresses a different need from supplying information for a task:

Choose the route that matches the need
Your needStarting routeExample
Use a known piece of information now.Supply it as context.Provide the approved order excerpt for one draft.
Find relevant information in a changing collection.Retrieve it from its source.Look up the current order each time a follow-up is prepared.
Improve a recurring pattern of model performance.Test instructions and examples first; consider fine-tuning for an evidenced remaining need.Evaluate whether training on suitable examples improves a specialised extraction task.

These routes can be combined. Fine-tuning for an extraction pattern does not remove the need to retrieve today's order.

Uploading a policy is not, by itself, fine-tuning. Updating a search index is not retraining the language model. A product may separately retain inputs or use them in a later training process under its terms, so distinguish what happens during this response from the provider's subsequent use of data.

What happens to information you provide

Consumer applications, business subscriptions and API services can have different terms and controls even when they use the same model. Check the actual product, account configuration and agreement. Training use, retention, human access and onward processing are separate questions. Not used for training does not mean no copy is retained or that authorised staff can never review it.

An application can also keep its own logs or send content through another service. Establish where the information goes and what each party keeps. OpenAI's API data-control documentation provides one concrete example, distinguishing training use, monitoring logs and application state.

Further reading: Google's embedding explanation and the original retrieval-augmented generation research.


Part 3: Permitted action

6. Tools and connectors

A tool is a capability the application makes available for use: searching records, performing a calculation, creating a document or sending a message. A tool call is a request to use that capability, normally including structured arguments such as a record identifier or date range.

The model can propose which tool to call and what values to supply. Software in the application, provider or connected service performs the operation and returns a result. That software must implement the checks on the request and its permissions. To establish whether I sent it is true, inspect the email service's result or the destination record, not just the model's sentence.

An API provides a software-to-software route for these requests. A connector packages a connection to a service, often handling its sign-in and available operations. It may expose several tools. A browser tool instead operates through a website interface. These routes can differ in reliability, coverage and permission handling.

  1. You ask: When is Northbank's current order due?
  2. The application describes its order-lookup tool and the information required to identify the correct order.
  3. The model requests that lookup using the order's identifier. If several orders match and the intended one is unclear, it must ask rather than choose silently.
  4. The application or connected service checks the request and permissions, then performs the lookup. In this successful example it returns Customer: Northbank. Confirmed delivery: Tuesday. A real lookup could instead return an error or an access denial.
  5. The returned record is supplied to the model, which answers The order record says Tuesday and identifies its source. The request to look up the order and the record returned by the service are different parts of this exchange.

That returned material can fill the context window too. Good tools return enough information to support the task without dumping an entire database into every response. Apply the same distinction between source content and instructions to tool results. A failed lookup should remain a failed lookup, not silently become an answer from model memory.

OpenAI's function-calling guide provides one concrete implementation of this request-and-result cycle.

A common language for connections

Model Context Protocol, or MCP, standardises how compatible AI applications communicate with services that expose capabilities such as tools and resources. An MCP server supplies capabilities; the application's client communicates with it. This can reduce the need to build a different integration for every application.

The standard does not make every system compatible with every other system, create an account for you or grant business authority. A connection still depends on supported features, credentials and permissions. A connector labelled "connected" proves less than successfully retrieving the expected record under the intended identity.

The official MCP architecture overview explains the parts and their responsibilities.

Skills, plugins and computer use

A skill packages reusable instructions and supporting material for a kind of work. It may include templates, references or scripts. In the Agent Skills format, a SKILL.md file could describe how to look up an order, prepare a follow-up and check that its draft was saved. The procedure tells the application how to use available capabilities; it does not supply a missing connection or retrain the model. See the Codex skill documentation for a product implementation.

A plugin is an extension package for an application. Depending on the platform, it may provide tools, connectors, skills or other features. Installing it and granting access are different steps. Names and packaging vary, so check what the extension actually supplies and what it is authorised to reach.

A computer-use agent interacts with an interface through capabilities such as viewing, clicking and typing. Browser agents can use page structure, visual observations or both. This can make work possible when a suitable API is unavailable, but page changes, login steps and ambiguous controls can interrupt it. Verify the resulting record or action rather than treating a sequence of clicks as proof of success.

For our follow-up, a CRM connector can supply account context. A document tool can save a draft. An email tool can deliver it. You can use the first two without granting the third. Grant the operations the current job needs, even if a connector offers more. This is how connections create useful options rather than forcing an all-or-nothing automation.


7. Identity, access and authority

A connector's consent screen may ask to read your contacts, see documents or send messages. Those are different capabilities, even if one button accepts them together. Before accepting, connect the requested access to the job: preparing a draft and sending it are different responsibilities.

Whose access would the connector use? Authentication establishes the identity making a request. Authorisation determines what that identity may do. Signing in establishes neither permission to see every company record nor permission to perform every action.

Credentials are the means used to establish or exercise access, such as keys or access tokens. Here, an access token means a credential, not a language-model token. Scopes describe permitted categories of access; record-level rules can narrow those permissions further. Consent may be one step in granting access, but a consent screen alone does not explain every downstream use.

With delegated access, an application acts on behalf of a user, constrained by both the permissions granted to the application and what that user is entitled to access. A broad application scope does not, by itself, give the user access to every record. With an application or service identity, it acts as a non-human system under its own assigned permissions. The right arrangement depends on the task. A personal assistant and a shared overnight process need not use the same identity.

Provider documentation: Microsoft's authentication and authorisation overview.

Access makes delegation possible

A sales assistant that can read the relevant customers and due commitments can save its owner time. It does not need access to unrelated payroll records. Permission to draft a follow-up need not include permission to send it, change a price or promise a delivery date.

Business authority and technical access must agree. You may tell the assistant it is allowed to issue a small refund, but that instruction does not give the payment service a valid identity or enforce a limit. Equally, an overpowered credential can technically permit actions the business never approved. The application must constrain the operation, not leave the entire boundary to interpretation by the model.

The practical starting point is to describe the capability you need and the actions you authorise. Complete personal login or consent through the service's supported interface when required. Keep credentials in the appropriate secure settings, rather than pasting them into ordinary chat or project notes. Check an expected read before considering a bounded test action.

Who acts when the follow-up runs next week and you are not at the keyboard? Name the identity configured for that run, whether delegated user access or a service identity. Check its access to the relevant customers and its separate authority to send. A schedule starts work; it does not supply a new identity or extend permission. If the required access has expired, report the limitation rather than silently borrowing a more powerful account.

Access also has a lifetime. Revocation withdraws it; expiry ends a credential's validity. Stopping the assistant is not the same as ending its token or session. An issued token may keep working until expiry or the service enforces revocation; withdrawing access does not erase copies already made. Microsoft's revocation guidance explains the timing and enforcement detail.

Try it without granting new access: read a connector's documented permissions or review an existing connection's settings. Identify whose account it uses, what it can read and what it can change. Which of those capabilities would the draft-only Northbank job actually need? Leave any new consent unapproved.


Part 4: Verified result

8. From a conversation to recurring work

A single model response ends. Continuing work requires software to decide what happens next. An agent commonly combines a model, tools and a task loop that can select further steps. The surrounding execution software is sometimes called a harness. It assembles context, manages calls and state, handles permissions and decides when to stop or ask for help. Products use these terms differently, so inspect the actual behaviour.

A workflow describes a process. Some steps can be fixed rules; others can use a model's judgement. An automation runs work in response to a trigger without a new instruction each time. Neither every workflow nor every automation needs an AI model.

State records where that process is: which items were handled, what remains pending and whether an action completed. A trigger starts it, perhaps on a schedule or after an event. A run that starts twice should not send two identical customer messages. A run that never starts cannot reliably report its own absence; another component must detect the missing completion.

Delegation and completion

One agent can request work from another when the platform supports that arrangement. Orchestration coordinates those assignments, dependencies and returns. Separate agents may have separate contexts and permissions. Sharing a computer or giving them job titles does not establish communication or access.

More agents also mean more coordination. They may repeat the same assumption or pass incomplete evidence between themselves. Use a specialist or a separate reviewer when it improves the result, not because every task needs an organisation chart. A straightforward calculation may need one tool call.

A useful completion record says what happened and how it was checked. "No commitments due", "customer system unavailable", "drafts prepared" and "messages delivered" are different outcomes. For the next follow-up run, retain which commitments were handled, where the drafts and supporting references are, any approval or delivery status, unresolved issues and the next trigger. The next run should not have to guess whether to repeat work or wait for a decision.

The project-state.md example from the memory section could hold this run state, or the application could keep it in a database. The useful property is that the next run can find the current record and distinguish completed work from an attempt. You can then delegate continuity to the process instead of restarting it from memory yourself.


9. Choosing and judging a useful system

Start with a job worth doing. Then ask what information, capabilities and evidence that job requires. This gives model choice, context management and integration a purpose. Buying more context does not resolve an undefined business objective; choosing the strongest model does not repair a connector returning the wrong customer's record.

An evaluation checks performance against defined expectations. For the follow-up process, useful checks might include an ordinary order, a changed commitment, conflicting records, unavailable data and a request outside the assistant's authority. Inspect the result and the route taken, not only whether the message sounds professional.

A benchmark reports performance on a particular collection of tasks under particular conditions. It can inform selection, but it does not establish suitability for your data, tools and approval process. Compare candidate configurations on representative work, including exceptions. Keep enough of the configuration and results to understand what changed when behaviour improves or deteriorates.

Precision, variation and agreement

Some tasks need exactness rather than plausible language. Use a calculator or tested code for consequential arithmetic and exact character counts. Retrieve and compare the original for a quotation, and open a web address rather than trusting a generated one. Rare, local or recently changed facts need suitable evidence. These are reasons to select the right working method, not claims that every model always fails at those tasks.

The same visible request can produce different answers. Sampling is one cause; changed context, tool results and the underlying service are others. A product name may stay the same while its model or routing changes. Repeating a request can reveal variation, but agreement between repeated answers is not independent confirmation that they are correct.

Sycophancy is behaviour that favours agreeing with or validating the user over maintaining an evidence-based answer. A model can abandon a correct conclusion when challenged or support an assumption embedded in the question. Ask what evidence supports or contradicts the claim, rather than testing it only with Are you sure? Research on sycophancy demonstrates this behaviour in evaluated models.

Automation bias is on the human side: giving an automated recommendation undue weight or failing to check it against contrary evidence. A fluent draft can make approval feel like a formality. Give the reviewer the relevant facts, unresolved issues and actual proposed action, with enough time and authority to disagree. A human click alone does not establish a meaningful review.

Check quality when work runs in parallel

A useful personal assistant may not yet be a dependable team service. Within authorised test capacity, compare representative jobs individually and at the team's expected demand. Keep inputs and settings fixed where possible; compare factual correctness, completion, response time and cost. Different wording is not automatically worse quality.

If quality falls, inspect failed lookups, timeouts, truncation, retries and fallback models before guessing why. Keep request identifiers and configuration details where available. The technical reference covers routing and batching, but these are not the only possible causes. This is a test of the complete service, whether hosted or self-operated.

Why didn't it do what I expected?

  • Learned capability: is this a job the chosen model handles well, or does it need a different method, such as a calculator for exact arithmetic? Compare the available models or modes on the actual task.
  • Available information: if it forgot a decision or invented a customer fact, did it receive the current record? If it stopped early, inspect context, output and task limits. If it was slow, separate reading, reasoning and generation from waits for tools or people.
  • Permitted action: if a lookup or save failed, did the tool exist and did the acting identity have the required access? More reasoning cannot repair an unavailable connection or grant permission.
  • Verified result: if it said the work was done, does the destination record or service result show completion? Distinguish a prepared draft, a failed attempt and a sent message.

You can move from AI is unreliable to the assistant had no current order status and the workflow let it substitute an answer. The second statement identifies something that can be investigated and changed.

The complete customer follow-up

Return to the request at the beginning: prepare a useful message, check the latest order, remember a preference and repeat next week. Here is how the parts fit together. This is a proposed design, not a claim that a system has been built:

  1. Begin with a draft. The model helps turn your intention into a message. Without the order record, a delivery date remains an unanswered question rather than a plausible detail to invent.
  2. Supply the current record. The order system provides the confirmed commitment. The draft uses that date and identifies its source. A conflicting warehouse note becomes a question for the responsible person, not a silent change to the promise.
  3. Carry forward a preference. Saved memory says you prefer short messages. It supplies the tone; the current order still supplies the date. Keep commitments and unresolved questions in the working record rather than relying only on a chat summary.
  4. Arrange a lookup and a place for the result. A tool can retrieve the current order and another can save the draft with its references. Test the returned record and saved document, including a missing-record case. This removes repeated copying, not the need for a source.
  5. Give those connections an appropriate identity. Before connecting real records, establish which customers it may read and where it may save. Draft-only authority does not include sending, changing prices or creating delivery commitments.
  6. Add a weekly trigger when authorised. Preserve completion state so a repeated trigger does not duplicate work. Return unresolved records to a person and arrange detection of a missed run. Next week's process retrieves current facts instead of recycling this week's answer.

A later version could send messages after you authorise that capability and establish the necessary review and delivery checks. Until then, completion means a checked draft is saved, not that Northbank has been contacted. The value is less repeated preparation and fewer missed commitments, with a clear account of what happened.

Make a short sketch of your system

Choose one job you want AI to help with. Describe its purpose, sources, tools, identity and what counts as done in ordinary language. Mark anything you do not yet know rather than inventing an answer. Here is a bounded first version of our follow-up system:

A customer-follow-up sketch

  • Purpose: prepare accurate follow-up drafts about outstanding customer commitments.
  • Sources: current CRM and order records. Unresolved conflicts return to the responsible record owner.
  • Tools: retrieve the permitted records and save drafts for review. Sending requires separate authority.
  • Identity: an approved user or service identity limited to the relevant customers, explicitly configured for any scheduled run.
  • Done: a checked draft and supporting references are saved in the agreed place, unresolved questions are flagged and no email has been sent.

The sketch makes the intended system discussable; it does not prove that the connections, permissions or checks work. Use the diagnostics above to identify what to inspect and verify. If the job later includes sending messages, its authority and completion evidence must change too.

You do not need to become a model engineer to make these distinctions. You need a clear purpose, a reasonable understanding of the system and evidence proportionate to the consequences. Begin with one useful result, establish how it is produced and expand when the next capability is justified.

Optional: have an assistant help you make the sketch

In a chat with an assistant that can read this guide, describe one job you want help with and paste the prompt unchanged. You can also start with the prompt and let it ask about the job. Use ordinary language and fictional or sanitised details; no connected accounts are needed. The result is a proposed system sketch and a first test, not an installed or approved system.

Help me sketch one useful AI-assisted job using the complete guide at https://mikkosniemela.com/ai-foundations-2026-10-01. If you cannot read it, say so and ask me for the relevant text; do not claim to have followed it.

Start from the job and context I have already described. If they are missing, ask what I want help with. Ask the next useful question in ordinary language, reusing my answers rather than making me fill in a technical template. Establish the intended result, relevant sources, necessary capabilities, identity and limits, and what would count as completion. Distinguish what I have confirmed from your proposals and anything unknown. Do not invent available tools, permissions or records. Use fictional or sanitised examples and do not request credentials or confidential records.

Return a concise, readable sketch here in the conversation under five labels: Purpose; Sources; Tools; Identity; Done. Follow it with the unresolved questions or decisions and the smallest read-only experiment that could test the idea. State what to inspect and what result would support or contradict the proposed approach. The sketch is a proposal, not evidence that the system works.

Do not connect accounts, change files, run commands, send messages or activate recurring work. Keep consequential decisions with me.

Optional deployment track

10. Where models run: hardware, speed and cost

Start with four separate checks: can you obtain it, can it fit, can it work fast enough, and can you operate it dependably? A successful download answers only the first.

Can you try this on the laptop you already own? Would a workstation help? Does the company need to operate a service, or simply buy access to one? These questions become easier when you separate the model, the machine and the service people use.

The deployment choices and basic size, memory and speed distinctions are useful whichever route you choose. Treat the detailed calculations and dated hardware comparisons as an optional reference when evaluating a setup; you do not need to become a hardware buyer to use this guide.

Choose who operates the system

With hosted inference, a provider operates the model infrastructure and you use its application or API. With self-hosted inference, you operate the model service yourself, on owned or rented equipment. Local inference commonly means running it on your own device. A company can self-host in a cloud data centre without owning the physical server.

  • Provider application or API: the provider runs the model; you pay through a subscription, usage charges or an agreement. An enterprise account may add organisational controls without moving the model into your building.
  • Dedicated managed deployment: a provider operates a deployment allocated to your organisation. Confirm what is actually dedicated, where it runs and who administers it. A private network connection is not the same as owning the model or hardware.
  • Company-operated service: your team runs supported weights on rented cloud GPUs or on-premises equipment. You choose more of the configuration and take responsibility for capacity, updates, security and recovery.
  • Personal laptop or workstation: a local runtime loads a compatible model for your work. It can be a useful experiment or tool without being ready to serve the entire company.

Enterprise-grade describes an operating standard, not a parameter count. Identity and access management, separation between users, monitoring, dependable capacity, support and recovery matter. A small model can sit inside a well-run enterprise service. A large model on an unattended desktop does not become one merely because colleagues can connect to it. Managed inference endpoints illustrate one of the options between a shared API and operating everything yourself.

Why you cannot simply download Fable or Astra

September 25, 2026 example: Claude Fable 5.1 and GPT-6 Astra are available through hosted products and APIs, not as publicly downloadable weights for self-hosting. You can use them from home or work; you cannot install their weights on your own server. A desktop app is the interface, not evidence that the model runs on that desktop.

Even if weights become available, running them requires compatible software, enough memory, sufficient processing speed and an operating budget. At larger scales, interconnects, power, cooling and specialist staff also matter. Training a frontier model, running one copy and serving a worldwide customer base are different investments. Public documentation does not establish a defensible self-hosting bill for these two models, so an exact figure would be guesswork.

Memory capacity, bandwidth and compute

A CPU, or central processing unit, performs general-purpose computation. A GPU, or graphics processing unit, performs many operations in parallel and is widely used for model workloads. Other accelerators also exist. The runtime is the software environment that runs the model; its inference engine loads the model and performs the calculations. Support for the model format and hardware matters as much as having a powerful-looking machine.

RAM is system working memory; VRAM is memory associated with a discrete GPU. Apple silicon Macs and some other systems use unified memory accessible to both CPU and GPU. A PC with 64 GB of RAM and an 8 GB graphics card does not have 72 GB of interchangeable fast GPU memory. Nor is all of a Mac's advertised unified memory free for one model: the operating system, other applications and runtime need space too.

  • Capacity: how much can be held in working memory. The model weights, context state and runtime buffers must fit in the chosen arrangement.
  • Bandwidth: how quickly data can move from memory to the processor. Two machines with the same memory capacity can generate tokens at very different speeds.
  • Compute: how quickly the processors perform the required mathematics. Peak advertised operations per second are not the same as tokens per second for your model.

Disk storage holds the download; it is not a replacement for fast working memory. Offloading moves parts of the workload between GPU memory, system memory or storage. It can make a larger model runnable while slowing it substantially. Multiple GPUs can share a model only with suitable software and communication; their memory is not automatically one large pool. See loading and offloading large models.

Finding a model you can run

Open-weight means the trained weights are available under a licence. It does not necessarily mean unrestricted commercial use, available training data or the same openness as a fully open-source project. Some highly capable models have downloadable weights; others do not. Frontier describes a level of capability, not a file format or licence.

Hugging Face's Model Hub is a collection of model repositories: places where publishers share files, revisions and documentation. A model card explains intended use, limitations and often evaluation results. Begin with the original publisher, check the licence and find a format your runtime supports. Community conversions can be useful, but a familiar model name is not proof of origin or quality.

LM Studio provides an application for finding, downloading and chatting with local models. Ollama provides an app, command-line tools and an API for running supported models. They are software around a model, not model families themselves. LM Studio can run GGUF model files with the llama.cpp engine, or models packaged for Apple's MLX framework on Apple silicon. These are implementation choices, not alternative measures of intelligence. Both local and remote arrangements exist; verify where the selected model actually runs.

What makes a model large

A model described as 8B has roughly eight billion parameters. Its architecture determines the count through choices such as the number and width of layers, vocabulary representations and expert components. Increasing a context setting or asking the model to think harder does not add learned parameters.

More parameters can support greater capability when combined with suitable training, but size is not a universal quality ranking. Training data, architecture, post-training and the task all matter. A small specialist model may handle classification or extraction well; a more capable model may be worth the cost for an ambiguous analysis. Neither size gives the model today's private customer records unless the system supplies them.

The precision describes how the numbers are represented. A 16-bit value takes two bytes; a four-bit value takes half a byte in the simplest arithmetic. Quantisation uses lower-precision representations to reduce requirements, with effects on quality and speed that need testing. It does not turn a 70B model into an 8B model. Distillation can train a smaller model using a larger model's outputs, but the result is a different model, not the original in a smaller download.

Raw weight storage in bytes ≈ parameter count × bits per parameter ÷ 8

Calculated raw weight storage, before operating overhead
Parameters16-bit4-bit
8 billion16 GB4 GB
32 billion64 GB16 GB
70 billion140 GB35 GB

These are decimal gigabytes and simplified calculations, not total memory requirements. Real formats include metadata and may mix precisions. For example, Ollama's Qwen3 8B Q4_K_M download is about 5.2 GB, not the idealised 4 GB. Q4_K_M names a particular quantisation scheme. The quantisation overview explains the wider choices.

One model-size distinction affects this choice: a dense model uses its main processing layers for each token, while a mixture-of-experts (MoE) model uses a selection of expert components. Lower active computation does not mean only those active weights need storage. Check total weights as well as active parameters; the technical reference explains the routing and load questions.

The context setting changes the hardware bill

The weights are only part of the bill. Many transformer engines keep a KV cache: intermediate information about earlier tokens used to generate later ones. For a conventional full-attention cache, its size grows approximately with the tokens retained and the number of concurrent sequences. Longer documents and several simultaneous users can therefore make a previously comfortable configuration run out of memory.

The model's supported context-window limit, the runtime's configured limit and the tokens actually processed are different quantities. Some engines reserve memory for the configured capacity; others grow the cache as needed. Increasing a slider may reserve more memory before you fill it. It cannot create reliable million-token capability in a model that does not support it. Ollama's context settings and cache strategies show implementation differences.

The same model, a longer conversation

Calculated examples, not measured peak usage: the table combines published Q4_K_M download sizes with an estimated 16-bit KV cache for one sequence. It excludes runtime buffers, loading peaks, the operating system and other applications. Here 8K means 8,192 tokens and 32K means 32,768.

Approximate weights plus context cache, in decimal GB
ModelAt 8KAt 32KWhat to consider
Qwen3 8B
5.2 GB download
6.4 GB10.0 GBA compatible 16 GB laptop can be a starting point at modest context. Longer context leaves less room for everything else.
Qwen3 32B
20 GB download
22.1 GB28.6 GB32 GB becomes tight as context grows; a compatible 48–64 GB setup gives more headroom. This is a capacity estimate, not a speed promise.

Calculated from the 8B configuration, 32B configuration and Ollama's 32B download. Quantised caches, different engines and additional users change the result. Check the runtime estimate and observed peak before choosing equipment.

What tokens per second actually tells you

The engine first processes the supplied input, often called prefill, then generates output, often called decoding. Processing a long prompt can involve substantial parallel computation. Generating one sequence often depends heavily on memory bandwidth, although context, architecture, batching and the engine can change the bottleneck. More memory capacity alone does not make either phase fast.

Time to first token measures the wait before output begins, including relevant queuing and processing. Generation tokens per second describes the rate after generation begins. Aggregate throughput adds work across users: 1,000 tokens per second across a server is not necessarily 1,000 for each person. With reasoning models, the wait for visible text can also include generated reasoning that the interface does not display. Compare the same measurement, model, context and concurrency. Tokenisers can differ too.

For scale, 500 output tokens at 10 tokens per second take roughly 50 seconds of generation; at 50 tokens per second, roughly 10 seconds. Both exclude the initial wait and tool work. These are illustrative calculations, not hardware benchmarks. NVIDIA documents the performance measurements and the different inference bottlenecks.

Published result: Hugging Face reports about 115 generation tokens per second for gpt-oss-20b in bfloat16 on an M3 Ultra Mac. That is the publisher's reported result, not a benchmark performed for this guide or a promise for a long prompt, another engine or several users. A downloadable model from a provider is not necessarily the same model offered in its flagship hosted product.

The dated cost examples show the scale: an experiment on a compatible laptop may require no new hardware, while a workstation can cost several thousand dollars and rented GPUs charge by provisioned time. For Northbank, connecting current records and authorising actions are separate decisions from buying a larger machine.

Optional first run: try a small local model

Begin with a harmless task on equipment you already have. You are testing the relationship between the model, context, machine and result, not committing to a company platform.

  1. Check the runtime's system requirements and install LM Studio or Ollama from its official site. Choose local execution rather than a cloud model or remote server.
  2. For a concrete starting example, select a compatible Qwen3 8B Q4_K_M model. In Ollama, ollama run qwen3:8b downloads and starts the library model. In LM Studio, search for the model and inspect the publisher, format and download size before loading it.
  3. Start with a modest supported context, such as 8,192 tokens, and check the memory estimate. Do not maximise the slider merely because it is available. LM Studio also provides a resource-estimation command; Ollama's ollama ps shows the loaded model, allocated context and CPU/GPU split.
  4. Give it a short public document, ask for a summary and check the facts against the original. Observe memory use, the initial wait and generation speed. Record the model, quantisation, runtime and context so the comparison means something.
  5. Increase one demand at a time: a longer document, a larger model or a second user. If memory pressure or waiting becomes unacceptable, reduce the demand or evaluate another deployment option. Unload the model when finished to release its working memory.

Keep the first experiment local and disconnected from company accounts. Running a local model does not automatically prevent a connector, plugin or cloud fallback from sending information elsewhere. For the customer-follow-up example, first compare draft quality using fictional or public facts; connect real business records only after assessing the information and action boundaries.


Technical and cost reference

This reference is for readers investigating a deployment or an observed performance problem. It preserves the detailed mechanisms, calculations and September 25, 2026 product and price examples separately from the main learning path. Prices are order-of-magnitude planning references, not current quotations.

Dense models and mixtures of experts

A dense model uses its main processing layers for each token. A mixture-of-experts, or MoE, model routes each token through a selection of expert components alongside shared parts. The experts are learned computational components, not separate agents with reliable labels such as accountant or lawyer.

For example, Qwen3-30B-A3B has about 30.5B total parameters and 3.3B active for a token. That can reduce arithmetic compared with activating a similarly sized dense model. It does not give it the storage requirement of a 3.3B model: the other experts must remain available, whether in accelerator memory, another device or an offloading arrangement. Routing and communication also affect performance. Hugging Face's MoE explanation develops this distinction.

Can a busy expert make answers worse?

An expert's learned capability is not used up by earlier requests. However, tokens from similar requests can concentrate work on particular experts: expert load imbalance. In some implementations, an expert has a token-capacity limit within a batch, a group processed together. Overflow can change which expert computations are performed; token dropping can mean bypassing an expert's contribution, not deleting a word from the user's message. That can affect output quality. This MoE explanation describes capacity limits and overflow.

It is not inevitable. The DeepSeek-V3 report describes a deployment without token dropping during inference. Batching can also change numerical behaviour in dense and MoE models; vLLM's batch-invariance documentation addresses that separate issue. A quality drop under load does not identify its cause or establish that dense models are generally better. Test the actual workload under its expected demand.

How the context-memory estimate works

For these full-attention models, approximate cache bytes = 2 × layers × KV heads × head dimension × bytes per cache value × retained tokens × concurrent sequences. The first factor accounts for keys and values. The 8B configuration uses 36 layers, eight KV heads and a head dimension of 128; the 32B uses 64 layers with the same KV-head count and dimension. At two bytes per value, their caches use about 147 KB and 262 KB per retained token respectively.

This formula is not universal. Sliding-window attention can bound some layers' history; other architectures compress or represent state differently. Cache quantisation is separate from weight quantisation, and can have its own quality and compatibility trade-offs. GB uses one billion bytes; GiB uses 1,073,741,824 bytes, so software and hardware labels may differ even before overhead.

From an existing laptop to company infrastructure

Price examples checked September 25, 2026, in USD before tax. These illustrate scale, not a shopping recommendation. Hardware figures are computer-only prices, not memory-upgrade prices; displays, peripherals, support and electricity can add cost. The model-fit observations are planning estimates, not tested configurations.

Different starting points, different responsibilities
Starting pointCost referenceWhat the capacity means
Existing compatible 16 GB laptop$0 additional hardware purchaseStart with a small quantised model and modest context, such as the 8B example above. CPU-only operation may be slower; integrated and discrete GPUs differ. Local processing still uses power and your machine's time.
Mac Studio M5 Max
36 GB unified memory
From $2,499More space for models in the tens of billions of parameters at reduced precision. The 32B example with short context is plausible to evaluate; increasing its context consumes the remaining headroom.
Mac Studio M5 Ultra
96 GB unified memory
From $5,499A 70B model at roughly four-bit weights is a plausible capacity class to investigate. It does not follow that every 70B model, full context or multi-user workload will run well.
NVIDIA DGX Spark
128 GB unified memory
$4,699 published MSRPA Linux AI workstation with a different accelerator and software stack. More memory than another machine does not establish higher token speed or compatibility with every deployment recipe.
Rented H100 SXM GPU instance
80 GB GPU memory
$4.29 per hour in the cited single-GPU offerA way to test GPU infrastructure without buying it. At 720 hours, the compute charge is about $3,089. You still operate the software and pay for provisioned time, not just useful answers.

Sources: LM Studio system requirements; Apple's base configurations and announced prices; NVIDIA's Spark specification and price notice; Lambda instance pricing. Availability, regional prices and configurations change.

For a company, a successful single-user test is the beginning of sizing, not its conclusion. Several requests may share model weights but need additional context state and processing capacity. Add concurrency, queuing, peak demand and failure recovery to the assessment. A dedicated managed endpoint or a hosted API may be more sensible than maintaining equipment that sits idle. Self-hosting can still be justified by control, predictable demand or a particular deployment requirement.

What does a million-token context cost?

A million-token window is not a standard memory module you can price once. The model must support it, the engine must implement it, and both memory and response time must be acceptable. Buying enough RAM is not proof that a laptop can use that context correctly or promptly.

Provider documentation: the Qwen2.5-1M instructions specify at least 120 GB of GPU memory for the 7B model and 320 GB for the 14B model when processing one-million-token sequences. Those requirements belong to its specified vLLM serving engine and NVIDIA CUDA software stack, not to every quantisation or runtime. They illustrate why even a relatively small model can require substantial infrastructure at extreme context.

Illustrative rental budget using the cited H100 SXM offers
Model and documented memoryCapacity to evaluateCompute rental
Qwen2.5 7B-1M
At least 120 GB
2 × 80 GB = 160 GB2 × $4.19 = $8.38/hour
About $84 for 10 hours
Qwen2.5 14B-1M
At least 320 GB
4 × 80 GB = 320 GB4 × $4.09 = $16.36/hour
About $164 for 10 hours

The prices are calculated from Lambda's two- and four-GPU offers, not a tested deployment or a quote for completing a particular document. The four-GPU example only meets the documented minimum: configuration and any extra workload can require more headroom. Continuous use for 720 hours would cost about $6,034 or $11,779 in compute respectively, before tax and any additional charges. Provisioned but idle time also costs money.

A 128 GB unified-memory desktop is not automatically equivalent to the first GPU deployment. The operating system shares its memory, and the published recipe targets a different execution stack. A larger Mac Studio may be worth evaluating with a supported implementation, but this guide does not establish a tested million-token Mac configuration or purchase price. There is no basis for recommending a laptop purchase from the token count alone.

A hosted alternative has a different bill. At the cited September 25 Astra Standard API rates, one million uncached input tokens cost $20 because the request exceeds its long-context pricing threshold. Output and relevant tool use cost extra; the available window must also accommodate generation. This is not the cost of owning the model, and it is not a quality-equivalent comparison with Qwen. It shows why paying for occasional use can differ radically from provisioning a server.


Each guide has its own job and can be read independently. Follow the connection relevant to what you want to do:

  • Building with AI Agents: turn an idea into working software through implementation and independent review.
  • AI Bots as Managers: establish shared context, responsibilities and recurring processes under human authority.
  • AI Security Primer: examine how information, influence and capabilities can lead to harm, and what safeguards and evidence are needed.
  • Cyber Exposure Primer: understand how information outside your control becomes useful to attackers and how to turn findings into action.

Sources and further reading

The main text links to explanations at the point they are useful. These additional primary references support further study. Provider documentation describes a particular implementation, not a rule for every AI system. Illustrative scenarios, calculated estimates, the measured token example and published results are identified where they appear. Hardware and price examples retain their stated September 25 reference date; this teaching revision does not reprice them.

Version history

2026-10-01 · First published edition. Introduces learned capability, available information, permitted action and verified result through the Northbank follow-up. Includes context snapshots, practical experiments, Markdown memory, reasoning and cost comparisons, a system sketch and optional sketch prompt, with deployment choices and a separate technical reference. Read the frozen 2026-10-01 edition.

Development before the first edition

The first review draft appeared on September 24, 2026. September 25 revisions added practical deployment, local tools, memory calculations, dated costs and stronger follow-up examples. Reader feedback and teaching experience shaped the October 1 four-part journey, clearer concept ordering and cumulative Northbank ending. These were working drafts, not separate frozen editions.

The conceptual core is intended to remain stable while examples and explanations improve. I aim to review product, capacity and price examples about twice a year, with material corrections sooner. Check each example's stated date: this review cadence does not guarantee that a price or product remains unchanged between editions. Published prices and calculated planning estimates remain labelled separately. Earlier published snapshots remain unchanged.

Last updated: October 1, 2026.

$ echo "EOF"