An Agentic Mixture of Experts for DevOps with Sunil Mallya - #708 – The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) – Podcast

Episódios

π0: A Foundation Model for Robotics with Sergey Levine - #719
18 fev· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Sergey Levine, associate professor at UC Berkeley and co-founder of Physical Intelligence, to discuss π0 (pi-zero), a general-purpose robotic foundation model. We dig into the model architecture, which pairs a vision language model (VLM) with a diffusion-based action expert, and the model training "recipe," emphasizing the roles of pre-training and post-training with a diverse mixture of real-world data to ensure robust and intelligent robot learning. We review the data collection approach, which uses human operators and teleoperation rigs, the potential of synthetic data and reinforcement learning in enhancing robotic capabilities, and much more. We also introduce the team’s new FAST tokenizer, which opens the door to a fully Transformer-based model and significant improvements in learning and generalization. Finally, we cover the open-sourcing of π0 and future directions for their research.

The complete show notes for this episode can be found at https://twimlai.com/go/719.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia - #718
10 fev· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today we’re joined by Victor Dibia, principal research software engineer at Microsoft Research, to explore the key trends and advancements in AI agents and multi-agent systems shaping 2025 and beyond. In this episode, we discuss the unique abilities that set AI agents apart from traditional software systems–reasoning, acting, communicating, and adapting. We also examine the rise of agentic foundation models, the emergence of interface agents like Claude with Computer Use and OpenAI Operator, the shift from simple task chains to complex workflows, and the growing range of enterprise use cases. Victor shares insights into emerging design patterns for autonomous multi-agent systems, including graph and message-driven architectures, the advantages of the “actor model” pattern as implemented in Microsoft’s AutoGen, and guidance on how users should approach the ”build vs. buy” decision when working with AI agent frameworks. We also address the challenges of evaluating end-to-end agent performance, the complexities of benchmarking agentic systems, and the implications of our reliance on LLMs as judges. Finally, we look ahead to the future of AI agents in 2025 and beyond, discuss emerging HCI challenges, their potential for impact on the workforce, and how they are poised to reshape fields like software engineering.

The complete show notes for this episode can be found at https://twimlai.com/go/718.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Estão a faltar episódios?

Clique aqui para atualizar o feed.
Speculative Decoding and Efficient LLM Inference with Chris Lott - #717
4 fev· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Chris Lott, senior director of engineering at Qualcomm AI Research to discuss accelerating large language model inference. We explore the challenges presented by the LLM encoding and decoding (aka generation) and how these interact with various hardware constraints such as FLOPS, memory footprint and memory bandwidth to limit key inference metrics such as time-to-first-token, tokens per second, and tokens per joule. We then dig into a variety of techniques that can be used to accelerate inference such as KV compression, quantization, pruning, speculative decoding, and leveraging small language models (SLMs). We also discuss future directions for enabling on-device agentic experiences such as parallel generation and software tools like Qualcomm AI Orchestrator.

The complete show notes for this episode can be found at https://twimlai.com/go/717.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Ensuring Privacy for Any LLM with Patricia Thaine - #716
28 jan· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Patricia Thaine, co-founder and CEO of Private AI to discuss techniques for ensuring privacy, data minimization, and compliance when using 3rd-party large language models (LLMs) and other AI services. We explore the risks of data leakage from LLMs and embeddings, the complexities of identifying and redacting personal information across various data flows, and the approach Private AI has taken to mitigate these risks. We also dig into the challenges of entity recognition in multimodal systems including OCR files, documents, images, and audio, and the importance of data quality and model accuracy. Additionally, Patricia shares insights on the limitations of data anonymization, the benefits of balancing real-world and synthetic data in model training and development, and the relationship between privacy and bias in AI. Finally, we touch on the evolving landscape of AI regulations like GDPR, CPRA, and the EU AI Act, and the future of privacy in artificial intelligence.

The complete show notes for this episode can be found at https://twimlai.com/go/716.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
AI Engineering Pitfalls with Chip Huyen - #715
21 jan· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Chip Huyen, independent researcher and writer to discuss her new book, “AI Engineering.” We dig into the definition of AI engineering, its key differences from traditional machine learning engineering, the common pitfalls encountered in engineering AI systems, and strategies to overcome them. We also explore how Chip defines AI agents, their current limitations and capabilities, and the critical role of effective planning and tool utilization in these systems. Additionally, Chip shares insights on the importance of evaluation in AI systems, highlighting the need for systematic processes, human oversight, and rigorous metrics and benchmarks. Finally, we touch on the impact of open-source models, the potential of synthetic data, and Chip’s predictions for the year ahead.

The complete show notes for this episode can be found at https://twimlai.com/go/715.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Evolving MLOps Platforms for Generative AI and Agents with Abhijit Bose - #714
13 jan· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Abhijit Bose, head of enterprise AI and ML platforms at Capital One to discuss the evolution of the company’s approach and insights on Generative AI and platform best practices. In this episode, we dig into the company’s platform-centric approach to AI, and how they’ve been evolving their existing MLOps and data platforms to support the new challenges and opportunities presented by generative AI workloads and AI agents. We explore their use of cloud-based infrastructure—in this case on AWS—to provide a foundation upon which they then layer open-source and proprietary services and tools. We cover their use of Llama 3 and open-weight models, their approach to fine-tuning, their observability tooling for Gen AI applications, their use of inference optimization techniques like quantization, and more. Finally, Abhijit shares the future of agentic workflows in the enterprise, the application of OpenAI o1-style reasoning in models, and the new roles and skillsets required in the evolving GenAI landscape.

The complete show notes for this episode can be found at https://twimlai.com/go/714.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Why Agents Are Stupid & What We Can Do About It with Dan Jeffries - #713
16 dez 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Dan Jeffries, founder and CEO of Kentauros AI to discuss the challenges currently faced by those developing advanced AI agents. We dig into how Dan defines agents and distinguishes them from other similar uses of LLM, explore various use cases for them, and dig into ways to create smarter agentic systems. Dan shared his “big brain, little brain, tool brain” approach to tackling real-world challenges in agents, the trade-offs in leveraging general-purpose vs. task-specific models, and his take on LLM reasoning. We also cover the way he thinks about model selection for agents, along with the need for new tools and platforms for deploying them. Finally, Dan emphasizes the importance of open source in advancing AI, shares the new products they’re working on, and explores the future directions in the agentic era.

The complete show notes for this episode can be found at https://twimlai.com/go/713.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Automated Reasoning to Prevent LLM Hallucination with Byron Cook - #712
9 dez 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Byron Cook, VP and distinguished scientist in the Automated Reasoning Group at AWS to dig into the underlying technology behind the newly announced Automated Reasoning Checks feature of Amazon Bedrock Guardrails. Automated Reasoning Checks uses mathematical proofs to help LLM users safeguard against hallucinations. We explore recent advancements in the field of automated reasoning, as well as some of the ways it is applied broadly, as well as across AWS, where it is used to enhance security, cryptography, virtualization, and more. We discuss how the new feature helps users to generate, refine, validate, and formalize policies, and how those policies can be deployed alongside LLM applications to ensure the accuracy of generated text. Finally, Byron also shares the benchmarks they’ve applied, the use of techniques like ‘constrained coding’ and ‘backtracking,’ and the future co-evolution of automated reasoning and generative AI.

The complete show notes for this episode can be found at https://twimlai.com/go/712.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
AI at the Edge: Qualcomm AI Research at NeurIPS 2024 with Arash Behboodi - #711
3 dez 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Arash Behboodi, director of engineering at Qualcomm AI Research to discuss the papers and workshops Qualcomm will be presenting at this year’s NeurIPS conference. We dig into the challenges and opportunities presented by differentiable simulation in wireless systems, the sciences, and beyond. We also explore recent work that ties conformal prediction to information theory, yielding a novel approach to incorporating uncertainty quantification directly into machine learning models. Finally, we review several papers enabling the efficient use of LoRA (Low-Rank Adaptation) on mobile devices (Hollowed Net, ShiRA, FouRA). Arash also previews the demos Qualcomm will be hosting at NeurIPS, including new video editing diffusion and 3D content generation models running on-device, Qualcomm's AI Hub, and more!

The complete show notes for this episode can be found at https://twimlai.com/go/711.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
AI for Network Management with Shirley Wu - #710
19 nov 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Shirley Wu, senior director of software engineering at Juniper Networks to discuss how machine learning and artificial intelligence are transforming network management. We explore various use cases where AI and ML are applied to enhance the quality, performance, and efficiency of networks across Juniper’s customers, including diagnosing cable degradation, proactive monitoring for coverage gaps, and real-time fault detection. We also dig into the complexities of integrating data science into networking, the trade-offs between traditional methods and ML-based solutions, the role of feature engineering and data in networking, the applicability of large language models, and Juniper’s approach to using smaller, specialized ML models to optimize speed, latency, and cost. Finally, Shirley shares some future directions for Juniper Mist such as proactive network testing and end-user self-service.

The complete show notes for this episode can be found at https://twimlai.com/go/710.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Why Your RAG System Is Broken, and How to Fix It with Jason Liu - #709
11 nov 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Jason Liu, freelance AI consultant, advisor, and creator of the Instructor library to discuss all things retrieval-augmented generation (RAG). We dig into the tactical and strategic challenges companies face with their RAG system, the different signs Jason looks for to identify looming problems, the issues he most commonly encounters, and the steps he takes to diagnose these issues. We also cover the significance of building out robust test datasets, data-driven experimentation, evaluation tools, and metrics for different use cases. We also touched on fine-tuning strategies for RAG systems, the effectiveness of different chunking strategies, the use of collaboration tools like Braintrust, and how future models will change the game. Lastly, we cover Jason’s interest in teaching others how to capitalize on their own AI experience via his AI consulting course.

The complete show notes for this episode can be found at https://twimlai.com/go/709.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
An Agentic Mixture of Experts for DevOps with Sunil Mallya - #708
4 nov 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today we're joined by Sunil Mallya, CTO and co-founder of Flip AI. We discuss Flip’s incident debugging system for DevOps, which was built using a custom mixture of experts (MoE) large language model (LLM) trained on a novel "CoMELT" observability dataset which combines traditional MELT data—metrics, events, logs, and traces—with code to efficiently identify root failure causes in complex software systems. We discuss the challenges of integrating time-series data with LLMs and their multi-decoder architecture designed for this purpose. Sunil describes their system's agent-based design, focusing on clear roles and boundaries to ensure reliability. We examine their "chaos gym," a reinforcement learning environment used for testing and improving the system's robustness. Finally, we discuss the practical considerations of deploying such a system at scale in diverse environments and much more.

The complete show notes for this episode can be found at https://twimlai.com/go/708.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Building AI Voice Agents with Scott Stephenson - #707
28 out 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Scott Stephenson, co-founder and CEO of Deepgram to discuss voice AI agents. We explore the importance of perception, understanding, and interaction and how these key components work together in building intelligent AI voice agents. We discuss the role of multimodal LLMs as well as speech-to-text and text-to-speech models in building AI voice agents, and dig into the benefits and limitations of text-based approaches to voice interactions. We dig into what’s required to deliver real-time voice interactions and the promise of closed-loop, continuously improving, federated learning agents. Finally, Scott shares practical applications of AI voice agents at Deepgram and provides an overview of their newly released agent toolkit.

The complete show notes for this episode can be found at https://twimlai.com/go/707.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Is Artificial Superintelligence Imminent? with Tim Rocktäschel - #706
21 out 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Tim Rocktäschel, senior staff research scientist at Google DeepMind, professor of Artificial Intelligence at University College London, and author of the recently published popular science book, “Artificial Intelligence: 10 Things You Should Know.” We dig into the attainability of artificial superintelligence and the path to achieving generalized superhuman capabilities across multiple domains. We discuss the importance of open-endedness in developing autonomous and self-improving systems, as well as the role of evolutionary approaches and algorithms. Additionally, we cover Tim’s recent research projects such as “Promptbreeder,” “Debating with More Persuasive LLMs Leads to More Truthful Answers,” and more.

The complete show notes for this episode can be found at https://twimlai.com/go/706.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
ML Models for Safety-Critical Systems with Lucas García - #705
14 out 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Lucas García, principal product manager for deep learning at MathWorks to discuss incorporating ML models into safety-critical systems. We begin by exploring the critical role of verification and validation (V&V) in these applications. We review the popular V-model for engineering critical systems and then dig into the “W” adaptation that’s been proposed for incorporating ML models. Next, we discuss the complexities of applying deep learning neural networks in safety-critical applications using the aviation industry as an example, and talk through the importance of factors such as data quality, model stability, robustness, interpretability, and accuracy. We also explore formal verification methods, abstract transformer layers, transformer-based architectures, and the application of various software testing techniques. Lucas also introduces the field of constrained deep learning and convex neural networks and its benefits and trade-offs.

The complete show notes for this episode can be found at https://twimlai.com/go/705.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
AI Agents: Substance or Snake Oil with Arvind Narayanan - #704
7 out 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Arvind Narayanan, professor of Computer Science at Princeton University to discuss his recent works, AI Agents That Matter and AI Snake Oil. In “AI Agents That Matter”, we explore the range of agentic behaviors, the challenges in benchmarking agents, and the ‘capability and reliability gap’, which creates risks when deploying AI agents in real-world applications. We also discuss the importance of verifiers as a technique for safeguarding agent behavior. We then dig into the AI Snake Oil book, which uncovers examples of problematic and overhyped claims in AI. Arvind shares various use cases of failed applications of AI, outlines a taxonomy of AI risks, and shares his insights on AI’s catastrophic risks. Additionally, we also touched on different approaches to LLM-based reasoning, his views on tech policy and regulation, and his work on CORE-Bench, a benchmark designed to measure AI agents' accuracy in computational reproducibility tasks.

The complete show notes for this episode can be found at https://twimlai.com/go/704.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
AI Agents for Data Analysis with Shreya Shankar - #703
30 set 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Shreya Shankar, a PhD student at UC Berkeley to discuss DocETL, a declarative system for building and optimizing LLM-powered data processing pipelines for large-scale and complex document analysis tasks. We explore how DocETL's optimizer architecture works, the intricacies of building agentic systems for data processing, the current landscape of benchmarks for data processing tasks, how these differ from reasoning-based benchmarks, and the need for robust evaluation methods for human-in-the-loop LLM workflows. Additionally, Shreya shares real-world applications of DocETL, the importance of effective validation prompts, and building robust and fault-tolerant agentic systems. Lastly, we cover the need for benchmarks tailored to LLM-powered data processing tasks and the future directions for DocETL.

The complete show notes for this episode can be found at https://twimlai.com/go/703.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Stealing Part of a Production Language Model with Nicholas Carlini - #702
23 set 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Nicholas Carlini, research scientist at Google DeepMind to discuss adversarial machine learning and model security, focusing on his 2024 ICML best paper winner, “Stealing part of a production language model.” We dig into this work, which demonstrated the ability to successfully steal the last layer of production language models including ChatGPT and PaLM-2. Nicholas shares the current landscape of AI security research in the age of LLMs, the implications of model stealing, ethical concerns surrounding model privacy, how the attack works, and the significance of the embedding layer in language models. We also discuss the remediation strategies implemented by OpenAI and Google, and the future directions in the field of AI security. Plus, we also cover his other ICML 2024 best paper, “Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining,” which questions the use and promotion of differential privacy in conjunction with pre-trained models.

The complete show notes for this episode can be found at https://twimlai.com/go/702.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Supercharging Developer Productivity with ChatGPT and Claude with Simon Willison - #701
16 set 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Simon Willison, independent researcher and creator of Datasette to discuss the many ways software developers and engineers can take advantage of large language models (LLMs) to boost their productivity. We dig into Simon’s own workflows and how he uses popular models like ChatGPT and Anthropic’s Claude to write and test hundreds of lines of code while out walking his dog. We review Simon’s favorite prompting and debugging techniques, his strategies for sidestepping the limitations of contemporary models, how he uses Claude’s Artifacts feature for rapid prototyping, his thoughts on the use and impact of vision models, the role he sees for open source models and local LLMs, and much more.

The complete show notes for this episode can be found at https://twimlai.com/go/701.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Automated Design of Agentic Systems with Shengran Hu - #700
2 set 2024· The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Today, we're joined by Shengran Hu, a PhD student at the University of British Columbia, to discuss Automated Design of Agentic Systems (ADAS), an approach focused on automatically creating agentic system designs. We explore the spectrum of agentic behaviors, the motivation for learning all aspects of agentic system design, the key components of the ADAS approach, and how it uses LLMs to design novel agent architectures in code. We also cover the iterative process of ADAS, its potential to shed light on the behavior of foundation models, the higher-level meta-behaviors that emerge in agentic systems, and how ADAS uncovers novel design patterns through emergent behaviors, particularly in complex tasks like the ARC challenge. Finally, we touch on the practical applications of ADAS and its potential use in system optimization for real-world tasks.

The complete show notes for this episode can be found at https://twimlai.com/go/700.
- Ouvir Ouvir novamente Continuar A reproduzir…
- Ouvir depois Ouvir depois
Mostrar mais

Episódios

π0: A Foundation Model for Robotics with Sergey Levine - #719

AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia - #718

Speculative Decoding and Efficient LLM Inference with Chris Lott - #717

Ensuring Privacy for Any LLM with Patricia Thaine - #716

AI Engineering Pitfalls with Chip Huyen - #715

Evolving MLOps Platforms for Generative AI and Agents with Abhijit Bose - #714

Why Agents Are Stupid & What We Can Do About It with Dan Jeffries - #713

Automated Reasoning to Prevent LLM Hallucination with Byron Cook - #712

AI at the Edge: Qualcomm AI Research at NeurIPS 2024 with Arash Behboodi - #711

AI for Network Management with Shirley Wu - #710

Why Your RAG System Is Broken, and How to Fix It with Jason Liu - #709

An Agentic Mixture of Experts for DevOps with Sunil Mallya - #708

Building AI Voice Agents with Scott Stephenson - #707

Is Artificial Superintelligence Imminent? with Tim Rocktäschel - #706

ML Models for Safety-Critical Systems with Lucas García - #705

AI Agents: Substance or Snake Oil with Arvind Narayanan - #704

AI Agents for Data Analysis with Shreya Shankar - #703

Stealing Part of a Production Language Model with Nicholas Carlini - #702

Supercharging Developer Productivity with ChatGPT and Claude with Simon Willison - #701

Automated Design of Agentic Systems with Shengran Hu - #700