Video summary
دوره آموزش مهندسی هوش مصنوعی (AI Engineering)
Main summary
Key takeaways
Course Goals and Approach
The presenter introduces a practical AI Engineering course that combines theory with Python coding. Its goal is to show how to integrate existing, pretrained models into applications—not how to train new models or focus primarily on machine learning research.
The course assumes familiarity with Python. As preparation, the presenter recommends a 12-hour Python course on the Neon Learn channel.
The presenter distinguishes AI Engineering, which focuses on building applications around existing models, from ML Engineering, which more often involves developing, improving, or researching models.
Planned Course Projects
The presenter previews nine projects:
- Build a chatbot connected to an LLM, with conversation history, streaming responses, token limits, and history summarization.
- Classify text—including sentiment and topic classification—using Natural Language Inference (NLI) and Hugging Face models.
- Create semantic search using embeddings.
- Build a RAG system for medical information and learn the basics of LangChain.
- Create an AI agent that can use tools such as weather services and internet search.
- Control model output formats, including structured JSON.
- Manage prompts with templates and preserve conversation history using checkpoints and a database.
- Use LangGraph to build a writer–evaluator workflow that repeatedly improves an article until it reaches a target score.
- Build a multi-agent system in which specialized agents collaborate under a supervisor.
Development Environment Setup
The presenter demonstrates preparing a Python workspace in Visual Studio Code:
- Install Python and a code editor. The demonstration uses VS Code, but other editors are also suitable.
- Create a project folder and open it in VS Code.
- Install useful extensions, including Python support and Code Runner, and configure code formatting to run on save.
- Configure Code Runner to execute code in the terminal so programs can receive user input.
- Use a virtual environment so each project has its own packages and package versions. The presenter prefers Pipenv and demonstrates installing it, creating an environment, and selecting its Python interpreter in VS Code.
- If setup or version problems occur, use an AI assistant to troubleshoot or ask for help in the comments.
Building a Chatbot Connected to an LLM
The first chatbot begins as a simple Python program that receives user input and returns a fixed response. It is then connected to an LLM through an API.
The presenter explains AI Engineering as adding pretrained models to applications to provide AI capabilities. LLMs can be accessed through commercial APIs, locally hosted models, or third-party providers. The demonstration uses Groq as a provider and the OpenAI-compatible Python SDK as the interface.
The API key should not be included in source code. The presenter stores it in an environment file and reads it through an environment variable.
Requests include a model name and a list of messages. Common message roles are:
- System: Instructions about how the model should behave.
- User: The user’s message.
- Assistant: The model’s response.
The presenter emphasizes using client libraries, or “wrappers,” rather than making raw API requests manually. Wrappers simplify requests, provide a more consistent interface, and are commonly used in AI Engineering.
LLM Concepts and Request Parameters
An LLM generates a response one token at a time. A token may be a whole word, part of a word, or punctuation; tokenization and token IDs vary by model.
Models may be text-only or multimodal. The presenter names several model families, including OpenAI, Claude, Gemini, Llama, Mistral, Qwen, Grok, Cohere Command, Gemma, Phi, and Falcon.
Important request settings include:
- Model: Which model to use.
- Temperature: How variable or creative the output is. Lower values are generally more predictable, while higher values allow more variation.
- Maximum tokens: A cap on generated output, which can also prematurely cut off an answer.
- Streaming: Whether to return pieces of the answer as they are generated rather than waiting for the complete response.
API usage can be billed by input and output tokens. Token counts and model performance are useful for managing cost and comparing models. Latency can be considered in terms of the time to the first output token and the time required to generate subsequent tokens.
Giving the Chatbot Conversation Memory
LLMs do not automatically remember earlier requests. Each call is separate unless the application sends relevant conversation history again.
The chatbot maintains a message list. After each turn, it adds the user’s message and the assistant’s response, then sends that list with the next request. This allows the model to interpret follow-up questions in context—for example, understanding that a question about “its population” refers to a city mentioned earlier.
The presenter refactors the chatbot into a class and adds reporting to inspect the conversation and token usage. This illustrates observability: developers should inspect what is sent to the model and what comes back, rather than treating the model as a black box.
Streaming Responses
To display responses as they are generated, the presenter:
- Enables streaming in the request.
- Iterates through the returned chunks and reads each chunk’s
deltacontent. - Prints each available piece immediately instead of waiting for the full answer.
- Accumulates the pieces into a complete response so it can still be added to the conversation history.
- Uses a short artificial delay to make the streaming effect easier to see.
Managing Long Conversations
A model’s context window limits how many input and output tokens it can process at once. Sending the full history on every turn can increase cost and eventually exceed that limit.
The presenter introduces two strategies:
- Sliding window: Retain recent messages and discard older ones. System instructions should be preserved separately so they are not accidentally removed.
- Summarization: Summarize older messages while keeping recent messages intact, allowing important context to survive without resending the entire conversation.
The presenter prototypes summarizing every few turns, then discusses why a token-based threshold may be more reliable than counting turns: messages can differ greatly in length.
During testing, some streamed replies appear empty. The presenter investigates and identifies a likely interaction involving a reasoning model, hidden reasoning tokens, streaming, and a maximum-token cap. The lesson is to troubleshoot systematically, inspect responses, and avoid assuming the provider is at fault.
Sentiment Analysis, Prompting, and NLI
Sentiment analysis classifies text as positive, negative, or neutral. The presenter first sketches an LLM-based approach, then explains that a smaller, specialized model is often more suitable and economical for this task.
Prompt engineering can guide output format and behavior. The presenter demonstrates few-shot prompting by providing examples of inputs and desired labels before asking the model to classify new text.
The lesson then introduces Natural Language Inference (NLI). Given a premise and a hypothesis, an NLI model predicts whether the hypothesis is:
- Entailed by the premise.
- Contradicted by the premise.
- Neutral—not established either way.
NLI can support tasks beyond sentiment, including zero-shot classification, fact checking, duplicate detection, and checking whether a passage answers a question. Zero-shot classification involves supplying candidate labels to a model without training a new classifier on a labeled dataset.
The subtitles end while discussing BERT and DistilBERT. BERT is described as an encoder model suited to understanding text and context, while DistilBERT is introduced as a smaller, distilled version. The discussion contrasts encoder models, which focus on understanding, with decoder models, which generate text.
Speakers and Sources Featured
- Speaker: Mohammad Seyed Aghaei, the presenter and instructor.
- Channel: Neon Learn, also transcribed as “Noon Learn” in the subtitles.
- Platforms, libraries, tools, and services: Python, Visual Studio Code, Code Runner, Pipenv,
python-dotenvand environment variables, Groq, the OpenAI-compatible API and SDK, OpenAI, Hugging Face, Transformers,tiktoken, LangChain, LangGraph, and LangSmith. - Models and model families: GPT/OpenAI models, Claude, Gemini, Llama, Mistral, Qwen, Grok, Cohere Command, Gemma, Phi, Falcon, BERT, and DistilBERT.
Rate this summary
Your feedback will help improve summaries.
Improve this summary
Reprocess with a stronger model when the summary feels incomplete or inaccurate.
Translate summary in another language
Ask questions to this video
Chat for follow-up questions, clarifications, and source-backed answers.