Conference Talk

Session focus

  • AI & ML & Bots
  • Automation
  • Data
  • Research
100%
Dev
0%
Design
20%
Marketing
60%
Business

Schedule

Date

May 30, 2024 5:00 pm

Duration

30 min

Location

Lucerna Cinema


Session details

All major generative AI models have been trained using data scraped from the web. Applications of large language models (LLMs) often extract web data to provide up-to-date context using Retrieval Augmented Generation (RAG). Unfortunately, reliably collecting online data at scale is challenging due to issues like blocking, dynamic content rendering, and the sheer volume of data. In this talk, Jan will explain how you can establish an efficient web data extraction pipeline, clean the HTML to circumvent the “garbage in, garbage out” problem, and demonstrate how to use this in an LLM application.

For questions and further discussion find the speaker in the Speaker’s Corner right after their talk.

Meet your presenter

This Site Uses Cookies

For processing purposes, your consent is required, which you express by selecting "Allow all." You can also customise your settings.