mirror of
https://github.com/Second-Hand-Friends/kleinanzeigen-bot.git
synced 2026-10-09 15:21:20 +02:00
fix: support redesigned ad page layout (Astro island fallback) (#1249)
## Problem Kleinanzeigen is gradually rolling out a redesigned ad page built with [Astro](https://astro.build/) islands and Tailwind CSS. When the bot's automated browser receives this new layout, the page's JavaScript does not fully hydrate, so the legacy DOM selectors the bot relies on are absent: - `.galleryimage-large` — image gallery container (not rendered) - `.iconlist-text` — seller name wrapper (not rendered) - `#viewad-extra-info` — creation date container (not rendered) Additionally, `window.BelenConf` returns `null` intermittently in the automated browser (likely a timing/race condition or Akamai bot protection), causing a `TypeError: 'NoneType' object is not subscriptable` on unguarded `belen_conf["universalAnalyticsOpts"]["dimensions"]` accesses. This affects **downloads** of ads served the new layout — the bot crashes or silently fails to extract images, creation date, and seller name. ## Root Cause 1. **BelenConf None**: Two code paths (line 664 for `l3_category_id`, line 836 for `ad_attributes`) accessed `belen_conf["universalAnalyticsOpts"]["dimensions"]` without a None guard, unlike line 655 which already had an `isinstance(belen_conf, dict)` check. 2. **Redesigned layout**: The new Astro-based page embeds all ad data in an `<astro-island>` element's `props` attribute as a JSON-encoded string. The DOM elements the bot queries via Selenium-style selectors don't exist until the island hydrates — which often doesn't happen in the automated browser context. ## Fix ### 1. BelenConf None guards Both unguarded `belen_conf["universalAnalyticsOpts"]["dimensions"]` accesses (lines 664 and 836) are now guarded with `isinstance(belen_conf, dict)` and chained `.get()` calls, matching the existing safe pattern at line 655. ### 2. Astro island data extraction (`_extract_island_props`) Added a new method that: - Queries all `<astro-island>` elements on the page - Finds the one whose `props` attribute contains ad-data keys (`imageDetails`, `formattedCreationDate`, `userDetails`) - HTML-unescapes the `props` attribute (uses `"` encoding) - Parses the JSON and unwraps the outer `data` → `[0, {...}]` envelope - Returns the inner ad data dict (keys: `formattedCreationDate`, `imageDetails`, `userDetails`, `price`, etc.) Props values in the Astro format are wrapped as `[type_index, value]` tuples. A helper `_unwrap_island_value()` strips this envelope. ### 3. Fallback cascades Three extraction points now cascade from legacy DOM selectors to island data: | Data point | Primary (legacy) | Secondary (redesigned DOM) | Tertiary (island JSON) | |---|---|---|---| | **Images** | `.galleryimage-large` container | — | `imageDetails.imageList[].xxLargeUrl` | | **Creation date** | `#viewad-extra-info > div:nth-child(1) > span:nth-child(2)` | `#viewad-extra-info span` | `formattedCreationDate` | | **Seller name** | `.iconlist-text > a/span` (within `#viewad-contact`) | `a[href*='/s-bestandsliste.html']` (within `#viewad-contact`) | `userDetails.contactName` | ### 4. Defensive error handling - Creation date extraction raises a clear `TimeoutError` if all three methods fail, instead of silently propagating a `NoneType` error downstream. - Image island fallback wraps each image download in the existing exception handler, so a single failed URL doesn't abort the batch. ## Testing Verified against ad ID 3491221676 ("Teichzaun Elektrozaungerät für Teichschutz"): - **Before**: `TypeError: 'NoneType' object is not subscriptable` at extract.py line 664, then `TimeoutError` on missing `.galleryimage-large` / `.iconlist-text` / `#viewad-extra-info` - **After**: Successfully downloads 2 images (438 KB + 216 KB JPGs) + YAML config, extracts seller name ("DanielP"), creation date ("22.08.2026"), price, and description Legacy-layout ads continue to work unchanged — the island fallback only activates when the primary DOM selectors time out. Unit tests (`tests/unit/test_extract.py`): all 150 tests pass, including updated tests for the restructured contact extraction and navigation flow. ### Pre-existing test failures (not introduced by this PR) The following tests fail on the clean upstream `main` branch as well and are unrelated to this PR: - `test_browser_profile_configuration` — requires a Chrome/Chromium binary installed on the runner; fails in environments without a browser. - `test_update_checker.py` (18 tests) — requires the `pytest-mock` plugin (provides the `mocker` fixture); errors at collection/setup time when the plugin is not installed. ## Compatibility - No breaking changes to existing functionality - No new dependencies - All changes confined to `extract.py` - Works with both the legacy and redesigned page layouts <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for extracting ads from Kleinanzeigen’s redesigned layout. * Improved handling of seller details, contact information, creation dates, categories, and special attributes. * Added image downloads from redesigned galleries and embedded page data. * Improved navigation when redirected ad content is delayed or unavailable. * **Bug Fixes** * Added fallback handling for missing or malformed page elements. * Improved resilience when downloading images and extracting ad information. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Jens <1742418+1cu@users.noreply.github.com>
This commit is contained in:
1 parent
0a505d8654
commit
9870c0a59a
4 files changed
+514
-54
No files matched your search
@@ -18,6 +18,7 @@ downloaded-ads
|
||||
|
||||
# python
|
||||
__pycache__
|
||||
__pypackages__
|
||||
/dist
|
||||
|
||||
# IntelliJ
|
||||
|
||||
@@ -15,14 +15,14 @@ import urllib.error as urllib_error
|
||||
import urllib.request as urllib_request
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Final
|
||||
from typing import Any, Final, cast
|
||||
|
||||
from kleinanzeigen_bot.model.ad_model import ContactPartial
|
||||
|
||||
from .model.ad_model import OPTION_NAME_BY_CARRIER_CODE, AdPartial, validate_condition_api_mapping
|
||||
from .model.config_model import AutoPriceReductionConfig, Config
|
||||
from .utils import dicts, files, i18n, loggers, misc, reflect
|
||||
from .utils.web_scraping_mixin import Browser, By, Element, WebScrapingMixin
|
||||
from .utils.web_scraping_mixin import Browser, By, WebScrapingMixin
|
||||
|
||||
__all__ = [
|
||||
"AdExtractor",
|
||||
@@ -34,6 +34,7 @@ _BREADCRUMB_MIN_DEPTH:Final[int] = 2
|
||||
BREADCRUMB_RE = re.compile(r"/c(\d+)")
|
||||
_MAX_FILENAME_COMPONENT_LENGTH:Final[int] = 255
|
||||
_DOWNLOAD_STEM_SUFFIX_BUDGET:Final[int] = len("__img9999.jpeg")
|
||||
_ISLAND_ENVELOPE_LENGTH:Final[int] = 2
|
||||
_STAGING_DIR_PREFIX:Final[str] = ".tmp-"
|
||||
_BACKUP_DIR_PREFIX:Final[str] = ".bak-"
|
||||
_RMTREE_RETRY_ATTEMPTS:Final[int] = 5
|
||||
@@ -456,29 +457,94 @@ class AdExtractor(WebScrapingMixin):
|
||||
with open(img_path, "wb") as f:
|
||||
shutil.copyfileobj(response, f)
|
||||
return str(img_path)
|
||||
except (urllib_error.URLError, urllib_error.HTTPError, OSError, shutil.Error) as e:
|
||||
# Narrow exception handling to expected network/filesystem errors
|
||||
except (urllib_error.URLError, urllib_error.HTTPError, ValueError, OSError, shutil.Error) as e:
|
||||
# ValueError: malformed URL string from island data
|
||||
# URLError/HTTPError: network or server errors
|
||||
# OSError/shutil.Error: filesystem errors
|
||||
LOG.warning("Failed to download image %s: %s", url, e)
|
||||
return None
|
||||
|
||||
async def _download_images_from_ad_page(self, directory:str, ad_file_stem:str) -> list[str]:
|
||||
def _extract_island_image_urls(self, island_props:dict[str, Any]) -> list[str]:
|
||||
"""Return usable image URLs from Astro island image data."""
|
||||
image_details = self._unwrap_island_value(island_props.get("imageDetails", {}))
|
||||
if not isinstance(image_details, dict):
|
||||
return []
|
||||
|
||||
image_list = self._unwrap_island_value(image_details.get("imageList", []))
|
||||
if not isinstance(image_list, list):
|
||||
return []
|
||||
|
||||
image_urls:list[str] = []
|
||||
for image_entry in image_list:
|
||||
image_data = self._unwrap_island_value(image_entry) if isinstance(image_entry, list) else image_entry
|
||||
if not isinstance(image_data, dict):
|
||||
continue
|
||||
image_url_raw = image_data.get("xxLargeUrl") or image_data.get("xLargeUrl") or image_data.get("largeUrl")
|
||||
image_url = self._unwrap_island_value(image_url_raw)
|
||||
if isinstance(image_url, str):
|
||||
image_urls.append(image_url)
|
||||
|
||||
return image_urls
|
||||
|
||||
async def _download_images_from_island(self, directory:str, ad_file_stem:str, island_props:dict[str, Any]) -> list[str]:
|
||||
"""Download the images embedded in Astro island data."""
|
||||
try:
|
||||
image_urls = self._extract_island_image_urls(island_props)
|
||||
if not image_urls:
|
||||
return []
|
||||
|
||||
image_paths:list[str] = []
|
||||
image_filename_prefix = f"{ad_file_stem}__img"
|
||||
loop = asyncio.get_running_loop()
|
||||
for image_number, image_url in enumerate(image_urls, start = 1):
|
||||
image_path = await loop.run_in_executor(
|
||||
None,
|
||||
self._download_and_save_image_sync,
|
||||
image_url,
|
||||
directory,
|
||||
image_filename_prefix,
|
||||
image_number,
|
||||
)
|
||||
if image_path:
|
||||
image_paths.append(Path(image_path).name)
|
||||
|
||||
LOG.info("Downloaded %s from Astro component data.", i18n.pluralize("image", len(image_paths)))
|
||||
return image_paths
|
||||
except Exception as error:
|
||||
LOG.warning("Astro component image fallback failed: %s", error)
|
||||
return []
|
||||
|
||||
async def _download_images_from_ad_page(self, directory:str, ad_file_stem:str, *, island_props:dict[str, Any] | None = None) -> list[str]:
|
||||
"""
|
||||
Downloads all images of an ad.
|
||||
|
||||
:param directory: the path of the directory created for this ad
|
||||
:param ad_file_stem: the rendered filename stem shared by the ad config and images
|
||||
:param island_props: optional Astro island props from the redesigned layout
|
||||
:return: the relative paths for all downloaded images
|
||||
"""
|
||||
|
||||
n_images:int
|
||||
img_paths = []
|
||||
try:
|
||||
# download all images from box
|
||||
# Try legacy layout first (galleryimage-large class)
|
||||
image_box = await self.web_probe(By.CLASS_NAME, "galleryimage-large")
|
||||
if image_box is None:
|
||||
raise TimeoutError("No image area found.")
|
||||
if image_box is not None:
|
||||
images = await self.web_find_all(By.CSS_SELECTOR, ".galleryimage-element[data-ix] > img", parent = image_box)
|
||||
else:
|
||||
images = []
|
||||
|
||||
if not images:
|
||||
# Redesigned layout: images use Tailwind-style classes.
|
||||
# Scope to the main article's image gallery to avoid picking up
|
||||
# thumbnails from "other ads" or recommendation sections.
|
||||
images = await self.web_find_all(
|
||||
By.CSS_SELECTOR,
|
||||
"article img[src*='img.kleinanzeigen.de'][class*='object-contain']",
|
||||
)
|
||||
if not images:
|
||||
raise TimeoutError("No image area found.")
|
||||
|
||||
images = await self.web_find_all(By.CSS_SELECTOR, ".galleryimage-element[data-ix] > img", parent = image_box)
|
||||
n_images = len(images)
|
||||
LOG.info("Found %s.", i18n.pluralize("image", n_images))
|
||||
|
||||
@@ -504,10 +570,26 @@ class AdExtractor(WebScrapingMixin):
|
||||
LOG.info("Downloaded %s.", i18n.pluralize("image", dl_counter))
|
||||
|
||||
except TimeoutError: # some ads do not require images
|
||||
LOG.warning("No image area found. Continuing without downloading images.")
|
||||
# Last-resort fallback: extract image URLs from the Astro island JSON
|
||||
if island_props:
|
||||
img_paths = await self._download_images_from_island(directory, ad_file_stem, island_props)
|
||||
if not img_paths:
|
||||
LOG.warning("No image area found. Continuing without downloading images.")
|
||||
|
||||
return img_paths
|
||||
|
||||
@staticmethod
|
||||
def _is_valid_creation_date(value:object) -> bool:
|
||||
"""Return whether value is a calendar date in the page's DD.MM.YYYY format."""
|
||||
if not isinstance(value, str):
|
||||
return False
|
||||
try:
|
||||
day, month, year = value.split(".")
|
||||
datetime.fromisoformat(f"{year}-{month}-{day} 00:00:00")
|
||||
except (ValueError, IndexError):
|
||||
return False
|
||||
return True
|
||||
|
||||
def extract_ad_id_from_ad_url(self, url:str) -> int:
|
||||
"""
|
||||
Extracts the ID of an ad, given by its reference link.
|
||||
@@ -607,11 +689,28 @@ class AdExtractor(WebScrapingMixin):
|
||||
await self.web_open(str(id_or_url)) # navigate to URL directly given
|
||||
await self.web_sleep()
|
||||
|
||||
# handle the case that invalid ad ID given
|
||||
# Handle invalid IDs before waiting for an ad title: the ``k0`` page
|
||||
# has no ad content and must retain the documented False result.
|
||||
if self.page.url.endswith("k0"):
|
||||
LOG.error("There is no ad under the given ID.")
|
||||
return False
|
||||
|
||||
# When navigating via the search page, the initial page load completes
|
||||
# on the search URL before the JS redirect to the ad page fires.
|
||||
# Wait for the ad page content to actually appear.
|
||||
if reflect.is_integer(id_or_url):
|
||||
LOG.debug("After search redirect, current URL: %s", self.page.url)
|
||||
try:
|
||||
await self.web_find(By.ID, "viewad-title", timeout = self.effective_timeout("page_load"))
|
||||
except TimeoutError:
|
||||
LOG.debug("viewad-title not found after search redirect. Current URL: %s", self.page.url)
|
||||
# Force-reload the page to ensure the ad page content is present
|
||||
await self.web_open(self.page.url, reload_if_already_open = True)
|
||||
await self.web_sleep()
|
||||
LOG.debug("After force-reload, current URL: %s", self.page.url)
|
||||
# Verify ad content is now present after the reload
|
||||
await self.web_find(By.ID, "viewad-title", timeout = self.effective_timeout("page_load"))
|
||||
|
||||
# close (warning) popup, if given
|
||||
popup = await self.web_probe(By.ID, "vap-ovrly-secure")
|
||||
if popup is not None:
|
||||
@@ -620,6 +719,50 @@ class AdExtractor(WebScrapingMixin):
|
||||
await self.web_sleep()
|
||||
return True
|
||||
|
||||
async def _extract_island_props(self) -> dict[str, Any]:
|
||||
"""Extract the props JSON from the Astro island on redesigned ad pages.
|
||||
|
||||
Kleinanzeigen's redesigned ad pages embed all ad data in an
|
||||
``<astro-island>`` element's ``props`` attribute as a JSON-encoded
|
||||
string. When the page's JavaScript hasn't fully hydrated (common in
|
||||
automated browsers), the DOM elements the bot normally queries may be
|
||||
absent, but this data is always present in the HTML.
|
||||
|
||||
Returns an empty dict if the island or its props cannot be parsed.
|
||||
"""
|
||||
try:
|
||||
raw = await self.web_execute("""(() => {
|
||||
// Find the astro-island that contains ad data (imageDetails, userDetails, etc.)
|
||||
const islands = document.querySelectorAll('astro-island');
|
||||
for (const island of islands) {
|
||||
const props = island.getAttribute('props') || '';
|
||||
if (props.includes('imageDetails') || props.includes('formattedCreationDate') || props.includes('userDetails')) {
|
||||
return props;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
})()""")
|
||||
if not raw or not isinstance(raw, str):
|
||||
return {}
|
||||
# The props attribute uses " encoded quotes
|
||||
decoded = html.unescape(raw)
|
||||
props = json.loads(decoded)
|
||||
# The ad data is nested under 'data' -> [0, {...}]
|
||||
data_val = props.get("data", {})
|
||||
if isinstance(data_val, list) and len(data_val) == _ISLAND_ENVELOPE_LENGTH:
|
||||
data_val = data_val[1]
|
||||
if isinstance(data_val, dict):
|
||||
return cast(dict[str, Any], data_val)
|
||||
return cast(dict[str, Any], props) if isinstance(props, dict) else {}
|
||||
except Exception:
|
||||
return {}
|
||||
|
||||
def _unwrap_island_value(self, val:Any) -> Any:
|
||||
"""Unwrap an Astro island prop value from its [type_index, value] envelope."""
|
||||
if isinstance(val, list) and len(val) == _ISLAND_ENVELOPE_LENGTH and isinstance(val[0], int):
|
||||
return val[1]
|
||||
return val
|
||||
|
||||
async def _extract_title_from_ad_page(self) -> str:
|
||||
"""
|
||||
Extracts the title from an ad page.
|
||||
@@ -669,6 +812,9 @@ class AdExtractor(WebScrapingMixin):
|
||||
"""
|
||||
info:dict[str, Any] = {"active": active_override if active_override is not None else True}
|
||||
|
||||
# Extract embedded Astro island data (redesigned layout fallback)
|
||||
island_props = await self._extract_island_props()
|
||||
|
||||
# Get BelenConf data which contains accurate ad_type information
|
||||
belen_conf = await self.web_execute("window.BelenConf")
|
||||
|
||||
@@ -684,7 +830,7 @@ class AdExtractor(WebScrapingMixin):
|
||||
# append subcategory and change e.g. category "161/172" to "161/172/lautsprecher_kopfhoerer"
|
||||
# take subcategory from third_category_name as key 'art_s' sometimes is a special attribute (e.g. gender for clothes)
|
||||
# the subcategory isn't really necessary, but when set, the appropriate special attribute gets preselected
|
||||
if third_category_id := belen_conf["universalAnalyticsOpts"]["dimensions"].get("l3_category_id"):
|
||||
if isinstance(belen_conf, dict) and (third_category_id := belen_conf.get("universalAnalyticsOpts", {}).get("dimensions", {}).get("l3_category_id")):
|
||||
info["category"] += f"/{third_category_id}"
|
||||
|
||||
info["title"] = title
|
||||
@@ -713,11 +859,33 @@ class AdExtractor(WebScrapingMixin):
|
||||
info["price"], info["price_type"] = await self._extract_pricing_info_from_ad_page()
|
||||
info["shipping_type"], info["shipping_costs"], info["shipping_options"] = await self._extract_shipping_info_from_ad_page()
|
||||
info["sell_directly"] = await self._extract_sell_directly_from_ad_page()
|
||||
info["images"] = await self._download_images_from_ad_page(directory, ad_file_stem)
|
||||
info["contact"] = await self._extract_contact_from_ad_page()
|
||||
info["images"] = await self._download_images_from_ad_page(directory, ad_file_stem, island_props = island_props)
|
||||
info["contact"] = await self._extract_contact_from_ad_page(island_props = island_props)
|
||||
info["id"] = ad_id
|
||||
|
||||
creation_date = await self.web_text(By.CSS_SELECTOR, DOWNLOAD_CREATION_DATE_SELECTOR)
|
||||
# Extraction of creation date — try the legacy selector first, fall back
|
||||
# to a broader query that works on the redesigned (Tailwind) layout,
|
||||
# then finally try the Astro island embedded JSON.
|
||||
creation_date:str | None = None
|
||||
try:
|
||||
legacy_date = await self.web_text(By.CSS_SELECTOR, DOWNLOAD_CREATION_DATE_SELECTOR)
|
||||
if self._is_valid_creation_date(legacy_date):
|
||||
creation_date = legacy_date
|
||||
except TimeoutError:
|
||||
pass # Legacy layout may omit the creation-date element.
|
||||
if not creation_date:
|
||||
try:
|
||||
redesigned_date = await self.web_text(By.CSS_SELECTOR, "#viewad-extra-info span")
|
||||
if self._is_valid_creation_date(redesigned_date):
|
||||
creation_date = redesigned_date
|
||||
except TimeoutError:
|
||||
pass # Redesigned layout may lack #viewad-extra-info; fall through to island fallback
|
||||
if not creation_date and island_props:
|
||||
island_date = self._unwrap_island_value(island_props.get("formattedCreationDate"))
|
||||
if self._is_valid_creation_date(island_date):
|
||||
creation_date = island_date
|
||||
if not creation_date:
|
||||
raise TimeoutError(_("Could not extract creation date from any selector or Astro component data."))
|
||||
|
||||
# convert creation date to ISO format
|
||||
created_parts = creation_date.split(".")
|
||||
@@ -849,7 +1017,7 @@ class AdExtractor(WebScrapingMixin):
|
||||
|
||||
return category
|
||||
|
||||
async def _extract_special_attributes_from_ad_page(self, belen_conf:dict[str, Any]) -> dict[str, str]:
|
||||
async def _extract_special_attributes_from_ad_page(self, belen_conf:dict[str, Any] | None) -> dict[str, str]:
|
||||
"""
|
||||
Extracts the special attributes from an ad page.
|
||||
If no items are available then special_attributes is empty
|
||||
@@ -858,7 +1026,11 @@ class AdExtractor(WebScrapingMixin):
|
||||
"""
|
||||
|
||||
# e.g. "art_s:lautsprecher_kopfhoerer|condition_s:like_new|versand_s:t"
|
||||
special_attributes_str = belen_conf["universalAnalyticsOpts"]["dimensions"].get("ad_attributes")
|
||||
if isinstance(belen_conf, dict):
|
||||
opts = belen_conf.get("universalAnalyticsOpts", {})
|
||||
special_attributes_str = opts.get("dimensions", {}).get("ad_attributes")
|
||||
else:
|
||||
special_attributes_str = None
|
||||
if not special_attributes_str:
|
||||
return await self._extract_special_attributes_from_dom()
|
||||
special_attributes = dict(item.split(":") for item in special_attributes_str.split("|") if ":" in item)
|
||||
@@ -1034,7 +1206,7 @@ class AdExtractor(WebScrapingMixin):
|
||||
LOG.debug("Could not determine sell_directly status: %s", e)
|
||||
return None
|
||||
|
||||
async def _extract_contact_from_ad_page(self) -> ContactPartial:
|
||||
async def _extract_contact_from_ad_page(self, *, island_props:dict[str, Any] | None = None) -> ContactPartial:
|
||||
"""
|
||||
Processes the address part involving street (optional), zip code + city, and phone number (optional).
|
||||
|
||||
@@ -1058,12 +1230,45 @@ class AdExtractor(WebScrapingMixin):
|
||||
contact["zipcode"] = zipcode # e.g. 19372
|
||||
contact["location"] = location # e.g. Mecklenburg-Vorpommern - Steinbeck
|
||||
|
||||
contact_person_element:Element = await self.web_find(By.ID, "viewad-contact")
|
||||
name_element = await self.web_find(By.CLASS_NAME, "iconlist-text", parent = contact_person_element)
|
||||
try:
|
||||
name = await self.web_text(By.TAG_NAME, "a", parent = name_element)
|
||||
except TimeoutError: # edge case: name without link
|
||||
name = await self.web_text(By.TAG_NAME, "span", parent = name_element)
|
||||
contact_person_element = await self.web_probe(By.ID, "viewad-contact")
|
||||
|
||||
if contact_person_element is not None:
|
||||
# Legacy layout: name is inside .iconlist-text > a/span
|
||||
name_element = await self.web_probe(By.CLASS_NAME, "iconlist-text", parent = contact_person_element)
|
||||
if name_element is not None:
|
||||
try:
|
||||
name = await self.web_text(By.TAG_NAME, "a", parent = name_element)
|
||||
except TimeoutError: # edge case: name without link
|
||||
name = await self.web_text(By.TAG_NAME, "span", parent = name_element)
|
||||
else:
|
||||
# Redesigned layout: seller name is in a link to /s-bestandsliste.html?userId=
|
||||
name_link = await self.web_probe(
|
||||
By.CSS_SELECTOR,
|
||||
"a[href*='/s-bestandsliste.html']",
|
||||
parent = contact_person_element,
|
||||
)
|
||||
if name_link is not None:
|
||||
name = await self.extract_visible_text(name_link)
|
||||
elif island_props:
|
||||
# Island fallback: extract contact name from embedded JSON
|
||||
user_details = self._unwrap_island_value(island_props.get("userDetails", {}))
|
||||
island_name = self._unwrap_island_value(user_details.get("contactName", "")) if isinstance(user_details, dict) else ""
|
||||
name = str(island_name) if island_name else ""
|
||||
if not name:
|
||||
LOG.warning("Could not extract seller name from contact area or Astro component data.")
|
||||
else:
|
||||
LOG.warning("Could not extract seller name from contact area.")
|
||||
name = ""
|
||||
elif island_props:
|
||||
# No viewad-contact container; try island fallback directly
|
||||
user_details = self._unwrap_island_value(island_props.get("userDetails", {}))
|
||||
island_name = self._unwrap_island_value(user_details.get("contactName", "")) if isinstance(user_details, dict) else ""
|
||||
name = str(island_name) if island_name else ""
|
||||
if not name:
|
||||
LOG.warning("Could not extract seller name from contact area or Astro component data.")
|
||||
else:
|
||||
LOG.warning("Could not extract seller name from contact area.")
|
||||
name = ""
|
||||
contact["name"] = name
|
||||
|
||||
if "street" not in contact:
|
||||
|
||||
@@ -624,6 +624,10 @@ kleinanzeigen_bot/extract.py:
|
||||
"Downloaded %s.": "%s heruntergeladen."
|
||||
"No image area found. Continuing without downloading images.": "Keine Bildbereiche gefunden. Fahre ohne Bilder-Download fort."
|
||||
|
||||
_download_images_from_island:
|
||||
"Downloaded %s from Astro component data.": "%s aus Astro-Komponentendaten heruntergeladen."
|
||||
"Astro component image fallback failed: %s": "Bild-Fallback aus Astro-Komponentendaten fehlgeschlagen: %s"
|
||||
|
||||
_log_download_name_truncation:
|
||||
"Download name truncated {id} placeholder: template='%s', max_length=%d, id='%s', rendered='%s'": "Download-Name hat {id} Platzhalter gekürzt: Vorlage='%s', max_length=%d, id='%s', gerendert='%s'"
|
||||
"Download name truncated {title} placeholder: template='%s', max_length=%d, title='%s', rendered='%s'": "Download-Name hat {title} Platzhalter gekürzt: Vorlage='%s', max_length=%d, titel='%s', gerendert='%s'"
|
||||
@@ -657,8 +661,13 @@ kleinanzeigen_bot/extract.py:
|
||||
"Using existing folder for ad %s at %s.": "Verwende bestehenden Ordner für Anzeige %s unter %s."
|
||||
"Could not remove staging directory %s: %s": "Konnte Staging-Verzeichnis %s nicht entfernen: %s"
|
||||
|
||||
_extract_ad_page_info:
|
||||
"Could not extract creation date from any selector or Astro component data.": "Erstellungsdatum konnte weder über Selektoren noch über Astro-Komponentendaten extrahiert werden."
|
||||
|
||||
_extract_contact_from_ad_page:
|
||||
"No street given in the contact.": "Keine Straße in den Kontaktdaten angegeben."
|
||||
"Could not extract seller name from contact area or Astro component data.": "Verkäufername konnte weder aus dem Kontaktbereich noch aus Astro-Komponentendaten extrahiert werden."
|
||||
"Could not extract seller name from contact area.": "Verkäufername konnte aus dem Kontaktbereich nicht extrahiert werden."
|
||||
|
||||
_extract_category_from_ad_page:
|
||||
"Breadcrumb container 'vap-brdcrmb' not found; cannot extract ad category: %s": "Breadcrumb-Container 'vap-brdcrmb' nicht gefunden; kann Anzeigenkategorie nicht extrahieren: %s"
|
||||
|
||||
+275
-30
@@ -90,6 +90,36 @@ class TestAdExtractorBasics:
|
||||
assert extractor.config == test_bot_config
|
||||
assert extractor.download_dir == Path("downloaded-ads")
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_island_props_unescapes_and_unwraps_ad_data(self, test_extractor:extract_module.AdExtractor) -> None:
|
||||
"""Extract usable ad data from an HTML-escaped Astro props attribute."""
|
||||
astro_props = (
|
||||
"{"data":[0,{"formattedCreationDate":[0,"22.08.2026"],"
|
||||
""userDetails":[0,{"contactName":[0,"DanielP"]}]}]}"
|
||||
)
|
||||
|
||||
with patch.object(test_extractor, "web_execute", new_callable = AsyncMock, return_value = astro_props):
|
||||
island_props = await test_extractor._extract_island_props()
|
||||
|
||||
assert island_props["formattedCreationDate"] == [0, "22.08.2026"]
|
||||
assert island_props["userDetails"] == [0, {"contactName": [0, "DanielP"]}]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_island_props_ignores_malformed_data(self, test_extractor:extract_module.AdExtractor) -> None:
|
||||
"""Ignore malformed island attributes so legacy extraction can continue."""
|
||||
with patch.object(test_extractor, "web_execute", new_callable = AsyncMock, return_value = "{not json"):
|
||||
assert await test_extractor._extract_island_props() == {}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_island_props_returns_unwrapped_props_for_unexpected_data_shape(self, test_extractor:extract_module.AdExtractor) -> None:
|
||||
"""Retain parseable props when their data envelope has an unexpected shape."""
|
||||
with patch.object(test_extractor, "web_execute", new_callable = AsyncMock, return_value = '{"data":"unexpected"}'):
|
||||
assert await test_extractor._extract_island_props() == {"data": "unexpected"}
|
||||
|
||||
def test_unwrap_island_value_keeps_non_enveloped_values(self, test_extractor:extract_module.AdExtractor) -> None:
|
||||
"""Keep ordinary values intact when they do not use Astro's tuple envelope."""
|
||||
assert test_extractor._unwrap_island_value("plain value") == "plain value"
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("url", "expected_id"),
|
||||
[
|
||||
@@ -459,6 +489,7 @@ class TestAdExtractorNavigation:
|
||||
patch.object(test_extractor, "page", page_mock),
|
||||
patch.object(test_extractor, "web_open", new_callable = AsyncMock) as mock_web_open,
|
||||
patch.object(test_extractor, "web_probe", new_callable = AsyncMock, return_value = MagicMock()),
|
||||
patch.object(test_extractor, "web_find", new_callable = AsyncMock, return_value = MagicMock()),
|
||||
patch.object(test_extractor, "web_click", new_callable = AsyncMock) as mock_web_click,
|
||||
):
|
||||
result = await test_extractor.navigate_to_ad_page(ad_id)
|
||||
@@ -495,19 +526,33 @@ class TestAdExtractorNavigation:
|
||||
page_mock = AsyncMock()
|
||||
page_mock.url = "https://www.kleinanzeigen.de/s-suchen.html?k0"
|
||||
|
||||
input_mock = AsyncMock()
|
||||
input_mock.clear_input = AsyncMock()
|
||||
input_mock.send_keys = AsyncMock()
|
||||
input_mock.apply = AsyncMock(return_value = True)
|
||||
input_mock.attrs = {}
|
||||
with (
|
||||
patch.object(test_extractor, "page", page_mock),
|
||||
patch.object(test_extractor, "web_open", new_callable = AsyncMock) as mock_web_open,
|
||||
patch.object(test_extractor, "web_find", new_callable = AsyncMock) as mock_web_find,
|
||||
):
|
||||
result = await test_extractor.navigate_to_ad_page(99999)
|
||||
|
||||
assert result is False
|
||||
mock_web_open.assert_awaited_once_with("https://www.kleinanzeigen.de/s-suchanfrage.html?keywords=99999")
|
||||
mock_web_find.assert_not_awaited()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_navigate_to_ad_page_reloads_when_search_redirect_has_no_ad_content(self, test_extractor:extract_module.AdExtractor) -> None:
|
||||
"""Retry once when the search page has not yet redirected to ad content."""
|
||||
page_mock = MagicMock()
|
||||
page_mock.url = "https://www.kleinanzeigen.de/s-anzeige/test/12345"
|
||||
|
||||
with (
|
||||
patch.object(test_extractor, "page", page_mock),
|
||||
patch.object(test_extractor, "web_open", new_callable = AsyncMock),
|
||||
patch.object(test_extractor, "web_find", new_callable = AsyncMock, return_value = input_mock),
|
||||
patch.object(test_extractor, "web_open", new_callable = AsyncMock) as mock_web_open,
|
||||
patch.object(test_extractor, "web_find", new_callable = AsyncMock, side_effect = [TimeoutError(), MagicMock()]),
|
||||
patch.object(test_extractor, "web_probe", new_callable = AsyncMock, return_value = None),
|
||||
):
|
||||
result = await test_extractor.navigate_to_ad_page(99999)
|
||||
assert result is False
|
||||
assert await test_extractor.navigate_to_ad_page(12345) is True
|
||||
|
||||
assert mock_web_open.await_args_list[1].args == (page_mock.url,)
|
||||
assert mock_web_open.await_args_list[1].kwargs == {"reload_if_already_open": True}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_own_ads_urls(self, test_extractor:extract_module.AdExtractor) -> None:
|
||||
@@ -999,6 +1044,78 @@ class TestAdExtractorContent:
|
||||
assert ad_cfg.created_on is not None
|
||||
assert ad_cfg.created_on.isoformat().startswith("2025-02-03")
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_ad_page_info_uses_island_creation_date_when_dom_is_missing(
|
||||
self,
|
||||
test_extractor:extract_module.AdExtractor,
|
||||
) -> None:
|
||||
"""Use the embedded date when neither supported DOM selector is available."""
|
||||
page_mock = MagicMock()
|
||||
page_mock.url = "https://www.kleinanzeigen.de/s-anzeige/test/12345"
|
||||
test_extractor.page = page_mock
|
||||
island_props = {"formattedCreationDate": [0, "22.08.2026"]}
|
||||
|
||||
with (
|
||||
patch.object(test_extractor, "_extract_island_props", new_callable = AsyncMock, return_value = island_props),
|
||||
patch.object(test_extractor, "web_execute", new_callable = AsyncMock, return_value = None),
|
||||
patch.object(
|
||||
test_extractor,
|
||||
"web_text",
|
||||
new_callable = AsyncMock,
|
||||
side_effect = ["Description text", TimeoutError(), TimeoutError()],
|
||||
),
|
||||
patch.multiple(
|
||||
test_extractor,
|
||||
_extract_category_from_ad_page = AsyncMock(return_value = "160"),
|
||||
_extract_special_attributes_from_ad_page = AsyncMock(return_value = {}),
|
||||
_extract_pricing_info_from_ad_page = AsyncMock(return_value = (None, "NOT_APPLICABLE")),
|
||||
_extract_shipping_info_from_ad_page = AsyncMock(return_value = ("NOT_APPLICABLE", None, None)),
|
||||
_extract_sell_directly_from_ad_page = AsyncMock(return_value = False),
|
||||
_download_images_from_ad_page = AsyncMock(return_value = []),
|
||||
_extract_contact_from_ad_page = AsyncMock(return_value = ContactPartial()),
|
||||
),
|
||||
):
|
||||
ad_cfg = await test_extractor._extract_ad_page_info("/some/dir", 12345, "ad_12345", "Test Title")
|
||||
|
||||
assert ad_cfg.created_on is not None
|
||||
assert ad_cfg.created_on.isoformat().startswith("2026-08-22")
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_ad_page_info_uses_island_creation_date_when_broader_selector_is_not_a_date(
|
||||
self,
|
||||
test_extractor:extract_module.AdExtractor,
|
||||
) -> None:
|
||||
"""Ignore unrelated extra-info text and retain the valid Astro creation date."""
|
||||
page_mock = MagicMock()
|
||||
page_mock.url = "https://www.kleinanzeigen.de/s-anzeige/test/12345"
|
||||
test_extractor.page = page_mock
|
||||
island_props = {"formattedCreationDate": [0, "22.08.2026"]}
|
||||
|
||||
with (
|
||||
patch.object(test_extractor, "_extract_island_props", new_callable = AsyncMock, return_value = island_props),
|
||||
patch.object(test_extractor, "web_execute", new_callable = AsyncMock, return_value = None),
|
||||
patch.object(
|
||||
test_extractor,
|
||||
"web_text",
|
||||
new_callable = AsyncMock,
|
||||
side_effect = ["Description text", TimeoutError(), "Anzeige online"],
|
||||
),
|
||||
patch.multiple(
|
||||
test_extractor,
|
||||
_extract_category_from_ad_page = AsyncMock(return_value = "160"),
|
||||
_extract_special_attributes_from_ad_page = AsyncMock(return_value = {}),
|
||||
_extract_pricing_info_from_ad_page = AsyncMock(return_value = (None, "NOT_APPLICABLE")),
|
||||
_extract_shipping_info_from_ad_page = AsyncMock(return_value = ("NOT_APPLICABLE", None, None)),
|
||||
_extract_sell_directly_from_ad_page = AsyncMock(return_value = False),
|
||||
_download_images_from_ad_page = AsyncMock(return_value = []),
|
||||
_extract_contact_from_ad_page = AsyncMock(return_value = ContactPartial()),
|
||||
),
|
||||
):
|
||||
ad_cfg = await test_extractor._extract_ad_page_info("/some/dir", 12345, "ad_12345", "Test Title")
|
||||
|
||||
assert ad_cfg.created_on is not None
|
||||
assert ad_cfg.created_on.isoformat().startswith("2026-08-22")
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_resolve_download_title_prefers_published_metadata_for_owned_overview(
|
||||
self,
|
||||
@@ -1460,6 +1577,14 @@ class TestAdExtractorCategory:
|
||||
|
||||
assert result == {}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_special_attributes_falls_back_when_belen_conf_is_missing(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Use the DOM fallback when BelenConf is unavailable."""
|
||||
with patch.object(extractor, "_extract_special_attributes_from_dom", new_callable = AsyncMock, return_value = {"condition_s": "ok"}):
|
||||
result = await extractor._extract_special_attributes_from_ad_page(None)
|
||||
|
||||
assert result == {"condition_s": "ok"}
|
||||
|
||||
@pytest.mark.asyncio
|
||||
# pylint: disable=protected-access
|
||||
async def test_extract_special_attributes_from_dom_skips_unrecognized_label(self, extractor:extract_module.AdExtractor) -> None:
|
||||
@@ -1536,8 +1661,8 @@ class TestAdExtractorContact:
|
||||
with (
|
||||
patch.object(extractor, "page", MagicMock()),
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock) as mock_web_text,
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock) as mock_web_find,
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [street_element, None]),
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [street_element, contact_person_element, name_element, None]),
|
||||
patch.object(extractor, "extract_visible_text", new_callable = AsyncMock, side_effect = visible_text_side_effect),
|
||||
):
|
||||
mock_web_text.side_effect = [
|
||||
@@ -1545,11 +1670,6 @@ class TestAdExtractorContact:
|
||||
"Test User",
|
||||
]
|
||||
|
||||
mock_web_find.side_effect = [
|
||||
contact_person_element,
|
||||
name_element,
|
||||
]
|
||||
|
||||
contact_info = await extractor._extract_contact_from_ad_page()
|
||||
assert contact_info.street == "Example Street 123"
|
||||
assert contact_info.zipcode == "12345"
|
||||
@@ -1569,6 +1689,66 @@ class TestAdExtractorContact:
|
||||
):
|
||||
await extractor._extract_contact_from_ad_page()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_contact_uses_island_name_without_legacy_contact_container(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Keep the seller name when the redesigned page omits viewad-contact."""
|
||||
island_props = {"userDetails": [0, {"contactName": [0, "DanielP"]}]}
|
||||
|
||||
with (
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock, return_value = "12345 Berlin - Mitte"),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, None, None]),
|
||||
):
|
||||
contact = await extractor._extract_contact_from_ad_page(island_props = island_props)
|
||||
|
||||
assert contact.name == "DanielP"
|
||||
assert contact.street is None
|
||||
assert contact.phone is None
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_contact_uses_redesigned_seller_link(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Use the redesigned seller-profile link when legacy name markup is absent."""
|
||||
contact_element = MagicMock()
|
||||
seller_link = MagicMock()
|
||||
|
||||
with (
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock, return_value = "12345 Berlin - Mitte"),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, contact_element, None, seller_link, None]),
|
||||
patch.object(extractor, "extract_visible_text", new_callable = AsyncMock, return_value = "DanielP"),
|
||||
):
|
||||
contact = await extractor._extract_contact_from_ad_page()
|
||||
|
||||
assert contact.name == "DanielP"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_contact_uses_span_when_legacy_name_has_no_link(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Support legacy seller names rendered without an anchor element."""
|
||||
contact_element = MagicMock()
|
||||
name_element = MagicMock()
|
||||
|
||||
with (
|
||||
patch.object(
|
||||
extractor,
|
||||
"web_text",
|
||||
new_callable = AsyncMock,
|
||||
side_effect = ["12345 Berlin - Mitte", TimeoutError(), "DanielP"],
|
||||
),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, contact_element, name_element, None]),
|
||||
):
|
||||
contact = await extractor._extract_contact_from_ad_page()
|
||||
|
||||
assert contact.name == "DanielP"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_contact_keeps_empty_name_when_no_seller_source_exists(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Return an empty seller name when neither DOM nor Astro data provides one."""
|
||||
with (
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock, return_value = "12345 Berlin - Mitte"),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, None, None]),
|
||||
):
|
||||
contact = await extractor._extract_contact_from_ad_page()
|
||||
|
||||
assert not contact.name
|
||||
|
||||
@pytest.mark.asyncio
|
||||
# pylint: disable=protected-access
|
||||
async def test_extract_contact_info_with_street_timeout(self, extractor:extract_module.AdExtractor) -> None:
|
||||
@@ -1580,12 +1760,11 @@ class TestAdExtractorContact:
|
||||
with (
|
||||
patch.object(extractor, "page", MagicMock()),
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock) as mock_web_text,
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock) as mock_web_find,
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [street_element, None]),
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [street_element, contact_person_element, name_element, None]),
|
||||
patch.object(extractor, "extract_visible_text", new_callable = AsyncMock, side_effect = TimeoutError()),
|
||||
):
|
||||
mock_web_text.side_effect = ["12345 Berlin - Mitte", "Test User"]
|
||||
mock_web_find.side_effect = [contact_person_element, name_element]
|
||||
|
||||
contact_info = await extractor._extract_contact_from_ad_page()
|
||||
assert contact_info.street is None
|
||||
@@ -1605,11 +1784,10 @@ class TestAdExtractorContact:
|
||||
with (
|
||||
patch.object(extractor, "page", MagicMock()),
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock) as mock_web_text,
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock) as mock_web_find,
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, phone_element]),
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, contact_person_element, name_element, phone_element]),
|
||||
):
|
||||
mock_web_text.side_effect = ["12345 Berlin - Mitte", "Test User", TimeoutError()]
|
||||
mock_web_find.side_effect = [contact_person_element, name_element]
|
||||
|
||||
contact_info = await extractor._extract_contact_from_ad_page()
|
||||
assert contact_info.street is None
|
||||
@@ -1629,16 +1807,11 @@ class TestAdExtractorContact:
|
||||
with (
|
||||
patch.object(extractor, "page", MagicMock()),
|
||||
patch.object(extractor, "web_text", new_callable = AsyncMock) as mock_web_text,
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock) as mock_web_find,
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, phone_element]),
|
||||
patch.object(extractor, "web_find", new_callable = AsyncMock),
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, side_effect = [None, contact_person_element, name_element, phone_element]),
|
||||
):
|
||||
mock_web_text.side_effect = ["12345 Berlin - Mitte", "Test User", "+49(0)1234 567890"]
|
||||
|
||||
mock_web_find.side_effect = [
|
||||
contact_person_element,
|
||||
name_element,
|
||||
]
|
||||
|
||||
contact_info = await extractor._extract_contact_from_ad_page()
|
||||
assert contact_info.phone == "01234567890" # Normalized phone number
|
||||
|
||||
@@ -1746,10 +1919,82 @@ class TestAdExtractorDownload:
|
||||
# pylint: disable=protected-access
|
||||
async def test_download_images_no_images(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Test image download when no images are found."""
|
||||
with patch.object(extractor, "web_probe", new_callable = AsyncMock, return_value = None):
|
||||
with (
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, return_value = None),
|
||||
patch.object(extractor, "web_find_all", new_callable = AsyncMock, return_value = []),
|
||||
):
|
||||
image_paths = await extractor._download_images_from_ad_page("/some/dir", "ad_12345")
|
||||
assert len(image_paths) == 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_download_images_uses_island_urls_after_missing_dom_gallery(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Download valid island image URLs and skip malformed entries independently."""
|
||||
island_props = {
|
||||
"imageDetails": [
|
||||
0,
|
||||
{
|
||||
"imageList": [
|
||||
0,
|
||||
[
|
||||
[0, {"xxLargeUrl": [0, "https://images.example/one.jpg"]}],
|
||||
[0, {"xLargeUrl": [0, "https://images.example/two.jpg"]}],
|
||||
[0, {"largeUrl": [0, None]}],
|
||||
],
|
||||
],
|
||||
},
|
||||
],
|
||||
}
|
||||
|
||||
with (
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, return_value = None),
|
||||
patch.object(extractor, "web_find_all", new_callable = AsyncMock, return_value = []),
|
||||
patch.object(
|
||||
extract_module.AdExtractor,
|
||||
"_download_and_save_image_sync",
|
||||
side_effect = ["/some/dir/ad_12345__img1.jpg", None],
|
||||
) as download_image,
|
||||
):
|
||||
image_paths = await extractor._download_images_from_ad_page("/some/dir", "ad_12345", island_props = island_props)
|
||||
|
||||
assert image_paths == ["ad_12345__img1.jpg"]
|
||||
assert download_image.call_args_list[0].args[0] == "https://images.example/one.jpg"
|
||||
assert download_image.call_args_list[1].args[0] == "https://images.example/two.jpg"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_download_images_uses_island_urls_when_legacy_gallery_is_empty(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Fall back when a legacy gallery container exists but contains no images."""
|
||||
island_props = {"imageDetails": [0, {"imageList": [0, [[0, {"xxLargeUrl": [0, "https://images.example/one.jpg"]}]]]}]}
|
||||
|
||||
with (
|
||||
patch.object(extractor, "web_probe", new_callable = AsyncMock, return_value = MagicMock()),
|
||||
patch.object(extractor, "web_find_all", new_callable = AsyncMock, return_value = []),
|
||||
patch.object(extract_module.AdExtractor, "_download_and_save_image_sync", return_value = "/some/dir/ad_12345__img1.jpg"),
|
||||
):
|
||||
image_paths = await extractor._download_images_from_ad_page("/some/dir", "ad_12345", island_props = island_props)
|
||||
|
||||
assert image_paths == ["ad_12345__img1.jpg"]
|
||||
|
||||
def test_extract_island_image_urls_skips_unusable_values(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Ignore malformed image structures and URLs from embedded Astro data."""
|
||||
assert extractor._extract_island_image_urls({"imageDetails": [0, []]}) == []
|
||||
assert extractor._extract_island_image_urls({"imageDetails": [0, {"imageList": [0, {}]}]}) == []
|
||||
malformed_entries = {
|
||||
"imageDetails": [
|
||||
0,
|
||||
{"imageList": [0, [[0, "invalid"], [0, {"largeUrl": [0, 42]}]]]},
|
||||
],
|
||||
}
|
||||
assert extractor._extract_island_image_urls(malformed_entries) == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_download_images_from_island_handles_missing_urls_and_extraction_errors(self, extractor:extract_module.AdExtractor) -> None:
|
||||
"""Treat unusable Astro image data as a non-fatal fallback failure."""
|
||||
empty_props = {"imageDetails": [0, {"imageList": [0, []]}]}
|
||||
assert await extractor._download_images_from_island("/some/dir", "ad_12345", empty_props) == []
|
||||
|
||||
with patch.object(extractor, "_extract_island_image_urls", side_effect = RuntimeError("bad data")):
|
||||
assert await extractor._download_images_from_island("/some/dir", "ad_12345", empty_props) == []
|
||||
|
||||
@pytest.mark.asyncio
|
||||
# pylint: disable=protected-access
|
||||
async def test_download_images_with_none_url(self, extractor:extract_module.AdExtractor) -> None:
|
||||
|
||||
Reference in new issue
Block a user