Skip to content

feat(core): add LowercaseTextProcessor, TrimWhitespaceProcessor, and LanguageDetectProcessor - #170

Open
Rajeev91691 wants to merge 1 commit into
google-gemini:mainfrom
Rajeev91691:feat/text-processors
Open

Rajeev91691 wants to merge 1 commit into
google-gemini:mainfrom
Rajeev91691:feat/text-processors

Conversation

@Rajeev91691

Copy link
Copy Markdown

This PR adds three new text-focused PartProcessors to genai_processors/core/text.py:

  1. LowercaseTextProcessor: A processor that lowercases incoming text parts.
  2. TrimWhitespaceProcessor: A processor that trims leading and trailing whitespace from incoming text parts.
  3. LanguageDetectProcessor: A processor that automatically detects the language of incoming text parts using langdetect and appends the language code to the part's metadata dictionary under a configurable key (defaulting to 'language').

It also adds complete unit tests for each of these in genai_processors/tests/text_test.py and updates pyproject.toml dependencies to include langdetect.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces three new text processors: LowercaseTextProcessor, TrimWhitespaceProcessor, and LanguageDetectProcessor (which utilizes the langdetect library), along with their corresponding unit tests and the addition of the langdetect dependency. The review feedback suggests initializing the DetectorFactory seed to ensure deterministic language detection, moving the inline langdetect import to the top of the file to adhere to style guidelines, and adding unit tests to cover edge cases such as fallback languages, custom metadata keys, and non-text parts.

Comment on lines +540 to +544
def __init__(
self, *, metadata_key: str = 'language', fallback_lang: str | None = None
):
self._metadata_key = metadata_key
self._fallback_lang = fallback_lang

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The langdetect library's detection algorithm is non-deterministic by default because it uses a randomized initialization. This can lead to inconsistent language detection results for the same input text across different runs, which can cause flaky tests or unpredictable behavior in production.

To ensure deterministic and consistent results, initialize the DetectorFactory seed to 0 during the processor's initialization.

Suggested change
def __init__(
self, *, metadata_key: str = 'language', fallback_lang: str | None = None
):
self._metadata_key = metadata_key
self._fallback_lang = fallback_lang
def __init__(
self, *, metadata_key: str = 'language', fallback_lang: str | None = None
):
self._metadata_key = metadata_key
self._fallback_lang = fallback_lang
from langdetect import DetectorFactory
DetectorFactory.seed = 0

async def call(
self, part: content_api.ProcessorPart
) -> AsyncIterable[content_api.ProcessorPartTypes]:
import langdetect

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since langdetect is listed as a required dependency in pyproject.toml, it is guaranteed to be installed. Importing it inline inside the call method on every invocation adds unnecessary overhead and is non-idiomatic.

It is highly recommended to move the import langdetect statement to the top of the file along with other module-level imports, adhering to the Google Python Style Guide.

References
  1. Imports should be placed at the top of the file, after any module comments and docstrings, and before module globals and constants. (link)

Comment on lines +780 to +782
results = await processor(input_part).gather()
self.assertEqual(len(results), 1)
self.assertEqual(results[0].metadata.get('language'), 'es')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To ensure the robustness of LanguageDetectProcessor, it is highly recommended to add unit tests covering the following edge cases:

  1. Fallback Language: Verifying that the fallback language is used when the input text is empty or contains no detectable language features (e.g., only numbers/punctuation).
  2. Custom Metadata Key: Verifying that the detected language is stored under a custom metadata key if provided.
  3. Non-Text Parts: Verifying that non-text parts (e.g., images) are ignored by the processor and passed through unchanged.
    results = await processor(input_part).gather()
    self.assertEqual(len(results), 1)
    self.assertEqual(results[0].metadata.get('language'), 'es')

  async def test_detect_fallback_on_empty_text(self):
    processor = text.LanguageDetectProcessor(fallback_lang='fr')
    input_part = content_api.ProcessorPart('', mimetype='text/plain')
    results = await processor(input_part).gather()
    self.assertEqual(len(results), 1)
    self.assertEqual(results[0].metadata.get('language'), 'fr')

  async def test_detect_custom_metadata_key(self):
    processor = text.LanguageDetectProcessor(metadata_key='detected_lang')
    input_part = content_api.ProcessorPart(
        'This is a simple english text.', mimetype='text/plain'
    )
    results = await processor(input_part).gather()
    self.assertEqual(len(results), 1)
    self.assertEqual(results[0].metadata.get('detected_lang'), 'en')

  async def test_detect_non_text_passed_through(self):
    processor = text.LanguageDetectProcessor()
    input_part = content_api.ProcessorPart(b'image_bytes', mimetype='image/png')
    results = await processor(input_part).gather()
    self.assertEqual(len(results), 1)
    self.assertEqual(results[0], input_part)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant