You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix: make demo e2e catalog queries resilient to transient failures
The generate-demos CI job fails ~45% of the time with jq exit status 5
on the ClusterCatalog Quickstart scenario. Several issues contribute:
- jq -s (slurp mode) buffers the entire operatorhubio FBC response in
memory before processing, risking system errors on large catalogs
- catalog content queries run exactly once with no retry, so any
transient port-forward or network hiccup fails the step immediately
- bash() does not attach stderr to ExitError, making failures opaque
- with CatalogdHA, kubectl port-forward to the service may connect to
a non-leader pod that returns 404 (empty local cache) and sticks
with it for the entire retry window
Remove jq slurp mode so each JSON object is processed in constant
memory, prefixing filters with 'objects' to skip non-object values in
the FBC stream. Wrap CatalogContainsSomePackages, PackageHasSomeChannels,
and PackageHasSomeBundles in waitFor for retry on transient errors. Add
curl --compressed to handle gzip-encoded responses and --fail with
pipefail to detect HTTP errors. Re-establish dead port-forwards via
liveness checks and reset port-forwards on query failure so retries
can reach the leader pod. Inject stderr into ExitError in bash() to
match k8sClient diagnostics. Log catalog query errors at V(0) so CI
timeout failures are diagnosable.
Co-Authored-By: Claude <noreply@anthropic.com>
0 commit comments