Last seen February 27, 2024

GitHub Security Vulnerability

With the following crawler configuration: ```python from bs4 import BeautifulSoup as Soup url = "https://example.com" loader = RecursiveUrlLoader( url=url, max_depth=2, extractor=lambda x: Soup(x, "html.parser").text ) docs = loader.load() ``` An attacker in control of the contents of `https://example.com` could place a malicious HTML file in there with links like "https://example.completely.different/my_file.html" and the crawler would proceed to download that file as well even though `prevent_outside=True`. https://github.com/langchain-ai/langchain/blob/bf0b3cc0b5ade1fb95a5b1b6fa260e99064c2e22/libs/community/langchain_community/document_loaders/recursive_url_loader.py#L51-L51 Resolved in https://github.com/langchain-ai/langchain/pull/15559

Technical Severity
Low severity
Lifecycle Status

RESOLVED

What Happened

With the following crawler configuration: ```python from bs4 import BeautifulSoup as Soup url = "https://example.com" loader = RecursiveUrlLoader( url=url, max_depth=2, extractor=lambda x: Soup(x, "html.parser").text ) docs = loader.load() ``` An attacker in control of the contents of `https://example.com` could place a malicious HTML file in there with links like "https://example.completely.different/my_file.html" and the crawler would proceed to download that file as well even though `prevent_outside=True`. https://github.com/langchain-ai/langchain/blob/bf0b3cc0b5ade1fb95a5b1b6fa260e99064c2e22/libs/community/langchain_community/document_loaders/recursive_url_loader.py#L51-L51 Resolved in https://github.com/langchain-ai/langchain/pull/15559

Why This Matters

Current evidence identifies a security issue involving GitHub, but does not yet support a more specific impact claim.

Recommended Action

No confirmed vendor remediation is available in the current evidence. Confirm whether GitHub is present in your environment and review the affected configuration.

Exposure

My Interests Exposure

Exposure unknown

Recommended Response
Last Seen

Feb 27, 2024 00:00

Exposure reason: This incident does not currently match a technology in My Interests.

Exploitation status: UNKNOWN

Primary entities:

GitHubvulnerabilitylangchain-ai/langchainlangchain

Timeline

  • Incident first seen
    Feb 27, 2024 00:00

    BugSkan first recorded this incident.

  • langchain Server-Side Request Forgery vulnerability
    Feb 27, 2024 00:00

    GitHub Advisory Database · Research

Sources

langchain Server-Side Request Forgery vulnerability

GitHub Advisory Database · Feb 27, 2024 00:00

With the following crawler configuration: ```python from bs4 import BeautifulSoup as Soup url = "https://example.com" loader = RecursiveUrlLoader( url=url, max_depth=2, extractor=lambda x: Soup(x, "html.parser").text ) docs = loader.load() ``` An attacker in control of the contents of `https://example.com` could place a malicious HTML file in there with links like "https://example.completely.different/my_file.html" and the crawler would proceed to download that file as well even though `prevent_outside=True`. https://github.com/langchain-ai/langchain/blob/bf0b3cc0b5ade1fb95a5b1b6fa260e99064c2e22/libs/community/langchain_community/document_loaders/recursive_url_loader.py#L51-L51 Resolved in https://github.com/langchain-ai/langchain/pull/15559

Open publisher source

My Interests Match

Want personalized relevance?

Create an account to see which incidents overlap with your interests.

← Back to incident intelligence