Repository navigation
feat(core): add LowercaseTextProcessor, TrimWhitespaceProcessor, and LanguageDetectProcessor - #170
Rajeev91691 wants to merge 1 commit into
Conversation
…LanguageDetectProcessor & unit tests
There was a problem hiding this comment.
Code Review
This pull request introduces three new text processors: LowercaseTextProcessor, TrimWhitespaceProcessor, and LanguageDetectProcessor (which utilizes the langdetect library), along with their corresponding unit tests and the addition of the langdetect dependency. The review feedback suggests initializing the DetectorFactory seed to ensure deterministic language detection, moving the inline langdetect import to the top of the file to adhere to style guidelines, and adding unit tests to cover edge cases such as fallback languages, custom metadata keys, and non-text parts.
| def __init__( | ||
| self, *, metadata_key: str = 'language', fallback_lang: str | None = None | ||
| ): | ||
| self._metadata_key = metadata_key | ||
| self._fallback_lang = fallback_lang |
There was a problem hiding this comment.
The langdetect library's detection algorithm is non-deterministic by default because it uses a randomized initialization. This can lead to inconsistent language detection results for the same input text across different runs, which can cause flaky tests or unpredictable behavior in production.
To ensure deterministic and consistent results, initialize the DetectorFactory seed to 0 during the processor's initialization.
| def __init__( | |
| self, *, metadata_key: str = 'language', fallback_lang: str | None = None | |
| ): | |
| self._metadata_key = metadata_key | |
| self._fallback_lang = fallback_lang | |
| def __init__( | |
| self, *, metadata_key: str = 'language', fallback_lang: str | None = None | |
| ): | |
| self._metadata_key = metadata_key | |
| self._fallback_lang = fallback_lang | |
| from langdetect import DetectorFactory | |
| DetectorFactory.seed = 0 |
| async def call( | ||
| self, part: content_api.ProcessorPart | ||
| ) -> AsyncIterable[content_api.ProcessorPartTypes]: | ||
| import langdetect |
There was a problem hiding this comment.
Since langdetect is listed as a required dependency in pyproject.toml, it is guaranteed to be installed. Importing it inline inside the call method on every invocation adds unnecessary overhead and is non-idiomatic.
It is highly recommended to move the import langdetect statement to the top of the file along with other module-level imports, adhering to the Google Python Style Guide.
References
- Imports should be placed at the top of the file, after any module comments and docstrings, and before module globals and constants. (link)
| results = await processor(input_part).gather() | ||
| self.assertEqual(len(results), 1) | ||
| self.assertEqual(results[0].metadata.get('language'), 'es') |
There was a problem hiding this comment.
To ensure the robustness of LanguageDetectProcessor, it is highly recommended to add unit tests covering the following edge cases:
- Fallback Language: Verifying that the fallback language is used when the input text is empty or contains no detectable language features (e.g., only numbers/punctuation).
- Custom Metadata Key: Verifying that the detected language is stored under a custom metadata key if provided.
- Non-Text Parts: Verifying that non-text parts (e.g., images) are ignored by the processor and passed through unchanged.
results = await processor(input_part).gather()
self.assertEqual(len(results), 1)
self.assertEqual(results[0].metadata.get('language'), 'es')
async def test_detect_fallback_on_empty_text(self):
processor = text.LanguageDetectProcessor(fallback_lang='fr')
input_part = content_api.ProcessorPart('', mimetype='text/plain')
results = await processor(input_part).gather()
self.assertEqual(len(results), 1)
self.assertEqual(results[0].metadata.get('language'), 'fr')
async def test_detect_custom_metadata_key(self):
processor = text.LanguageDetectProcessor(metadata_key='detected_lang')
input_part = content_api.ProcessorPart(
'This is a simple english text.', mimetype='text/plain'
)
results = await processor(input_part).gather()
self.assertEqual(len(results), 1)
self.assertEqual(results[0].metadata.get('detected_lang'), 'en')
async def test_detect_non_text_passed_through(self):
processor = text.LanguageDetectProcessor()
input_part = content_api.ProcessorPart(b'image_bytes', mimetype='image/png')
results = await processor(input_part).gather()
self.assertEqual(len(results), 1)
self.assertEqual(results[0], input_part)
This PR adds three new text-focused
PartProcessors togenai_processors/core/text.py:langdetectand appends the language code to the part'smetadatadictionary under a configurable key (defaulting to'language').It also adds complete unit tests for each of these in
genai_processors/tests/text_test.pyand updatespyproject.tomldependencies to includelangdetect.