Repository files navigation

Welcome to LaVague

Redefining internet surfing by transforming natural language instructions into seamless browser interactions.

🏄‍♀️ See LaVague in Action

Here are examples to show how LaVague can execute natural instructions on a browser to automate interactions with a website:

LaVague interacting with Hugging Face's website.

LaVague interacting with the IRS's website.

🎯 Motivations

LaVague is designed to automate menial tasks on behalf of its users. Many of these tasks are repetitive, time-consuming, and require little to no cognitive effort. By automating these tasks, LaVague aims to free up time for more meaningful endeavors, allowing users to focus on what truly matters to them.

By providing an engine turning natural language queries into Selenium code, LaVague is designed to make it easy for users or other AIs to automate easily express web workflows and execute them on a browser.

One of the key usages we see is to automate tasks that are personal to users and require them to be logged in, for instance automating the process of paying bills, filling out forms or pulling data from specific websites.

LaVague is built on open-source projects and leverages open-sources models, either locally or remote, to ensure the transparency of the agent and ensures that it is aligned with users' interests.

✨ Features

Natural Language Processing: Understands instructions in natural language to perform browser interactions.
Selenium Integration: Seamlessly integrates with Selenium for automating web browsers.
Open-Source: Built on open-source projects such as transformers and llama-index, and leverages open-source models, either locally or remote, to ensure the transparency of the agent and ensures that it is aligned with users' interests.
Local models for privacy and control: Supports local models like Gemma-7b so that users can fully control their AI assistant and have privacy guarantees.
Advanced AI techniques: Uses a local embedding (bge-small-en-v1.5) first to perform RAG to extract the most relevant HTML pieces to feed the LLM answering the query, as directly dropping the full HTML code would not fit in context. Then leverages Few-shot learning and Chain of Thought to elicit the most relevant Selenium code to perform the action without having to finetune the LLM (Nous-Hermes-2-Mixtral-8x7B-DPO) for code generation.

🚀 Getting Started

You can try LaVague in the following Colab notebook:

🗺️ Roadmap

This is an early project but could grow to democratize transparent and aligned AI models to undertake actions for the sake of users on the internet.

We see the following key areas to explore:

Fine-tune local models like a gemma-7b-it to be expert in Text2Action
Improve retrieval to make sure only relevant pieces of code are used for code generation
Support other browser engines (playwright) or even other automation frameworks

Keep up to date with our project backlog here.

🙋 Contributing

We would love your help in making La Vague a reality.

Please check out our contributing guide to see how you can get involved!

Please also join our Discord community where we can chat about the project further!

LaVague: Open-source Large Action Model to automate Selenium browsing

Repository files navigation

Welcome to LaVague

🏄‍♀️ See LaVague in Action

🎯 Motivations

✨ Features

🚀 Getting Started

🗺️ Roadmap

🙋 Contributing

Recommend

英特尔深耕零售业生态版图，发掘未来新机遇-品玩

AI Agents by B2B Rocket

厚积薄发40余载，康普综合布线解决方案Propel打造面向未来的智算中心

让李想付出代价的，是他心中的马斯克

达摩院牵头成立"无剑联盟"，探索RISC-V产业合作新范式 | 量子位

社群推广所用的短链接是如何生成的？

Spreadsheets are all you need

江苏半导体产业冲刺IPO，冰火两重天

IceBathList.com

无边框、纯透明！三星首款透明MICRO OLED电视亮相AWE 2024

About Joyk