Unify precomputation of aggregations behind a common API #16733

msfroh · 2024-11-27T22:40:08Z

Description

We've had a series of aggregation speedups that use the same strategy: instead of iterating through documents that match the query one-by-one, we can look at a Lucene segment and compute the aggregation directly (if some particular conditions are met).

In every case, we've hooked that into custom logic that hijacks the getLeafCollector method and throws CollectionTerminatedException. This creates the illusion that we're implementing a custom LeafCollector, when really we're not collecting at all (which is the whole point).

With this refactoring, the mechanism (hijacking getLeafCollector) is moved into AggregatorBase. Aggregators that have a strategy to precompute their answer can override tryPrecomputeAggregationForLeaf, which is expected to return true if they managed to precompute.

This should also make it easier to keep track of which aggregations have precomputation approaches (since they override this method).

Related Issues

N/A

Check List

~~Functionality includes testing.~~
~~API changes companion pull request created, if applicable.~~
~~Public documentation issue/PR created, if applicable.~~

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

We've had a series of aggregation speedups that use the same strategy: instead of iterating through documents that match the query one-by-one, we can look at a Lucene segment and compute the aggregation directly (if some particular conditions are met). In every case, we've hooked that into custom logic hijacks the getLeafCollector method and throws CollectionTerminatedException. This creates the illusion that we're implementing a custom LeafCollector, when really we're not collecting at all (which is the whole point). With this refactoring, the mechanism (hijacking getLeafCollector) is moved into AggregatorBase. Aggregators that have a strategy to precompute their answer can override tryPrecomputeAggregationForLeaf, which is expected to return true if they managed to precompute. This should also make it easier to keep track of which aggregations have precomputation approaches (since they override this method). Signed-off-by: Michael Froh <[email protected]>

github-actions · 2024-11-27T23:51:26Z

❌ Gradle check result for 4d5c32b: FAILURE

Please examine the workflow log, locate, and copy-paste the failure(s) below, then iterate to green. Is the failure a flaky test unrelated to your change?

bharath-techie · 2024-12-03T06:32:08Z

server/src/main/java/org/opensearch/search/aggregations/AggregatorBase.java

@@ -216,6 +220,10 @@ protected void preGetSubLeafCollectors(LeafReaderContext ctx) throws IOException
     */
    protected void doPreCollection() throws IOException {}

+    protected boolean tryPrecomputeAggregationForLeaf(LeafReaderContext ctx) throws IOException {


Are we planning to add support for subAggregations later ?

Don't we need something similar to LeafBucketCollector [ mainly the collect equivalent method ] - where the subAggregators can return the implementation and the parent class can invoke the subAggs implementation to collect each of the unique values/buckets.

[ This requirement is mainly for star tree aggregations ]

For example : for Terms aggs with sum sub-aggs, Terms | sum 1 500 --> Bucket 1 2 6000 --> Bucket 2

I was thinking that it would be the responsibility of an Aggregator to delegate to its subaggregators.

I think @sandeshkr419 managed to do that in #16674

sandeshkr419 · 2024-12-03T22:13:04Z

Regarding implementation of this, I have one more alternative which I think is worth discussing. How about bringing this abstraction at ContextIndexSearcher itself.

            weight = wrapWeight(weight);
            // See please https://github.com/apache/lucene/pull/964
            collector.setWeight(weight);
            leafCollector = collector.getLeafCollector(ctx);

Basically if we have pre computed aggregations already, we assign it as EarlyTerminationCollector.

So, what I'm thinking about is cases with sub-aggregations that we can pre-compute, which is highly relevant in cases of star tree pre-computation. For eg.: #16674 and if a dedicated abstraction for star-tree preCompute in ComtextIndexSearcher wopuld make more sense or not.

github-actions · 2024-12-12T20:50:38Z

✅ Gradle check result for 4d5c32b: SUCCESS

codecov · 2024-12-12T20:51:20Z

Codecov Report

Attention: Patch coverage is 97.10145% with 2 lines in your changes missing coverage. Please review.

Project coverage is 72.05%. Comparing base (10873f1) to head (4d5c32b).
Report is 53 commits behind head on main.

Files with missing lines	Patch %	Lines
...ket/terms/GlobalOrdinalsStringTermsAggregator.java	89.47%	1 Missing and 1 partial ⚠️

Additional details and impacted files

@@             Coverage Diff              @@
##               main   #16733      +/-   ##
============================================
- Coverage     72.15%   72.05%   -0.10%     
- Complexity    65145    65187      +42     
============================================
  Files          5315     5318       +3     
  Lines        303573   303836     +263     
  Branches      43925    43957      +32     
============================================
- Hits         219039   218937     -102     
- Misses        66587    67010     +423     
+ Partials      17947    17889      -58

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

msfroh added the skip-changelog label Nov 27, 2024

opensearch-ci-bot mentioned this pull request Nov 28, 2024

[AUTOCUT] Gradle Check Flaky Test Report for RemoteRestoreSnapshotIT #14324

Open

bharath-techie reviewed Dec 3, 2024

View reviewed changes

getsaurabh02 assigned msfroh Dec 9, 2024

msfroh marked this pull request as ready for review December 12, 2024 19:49

msfroh requested review from anasalkouz, andrross, ashking94, Bukhtawar, CEHENKLE, dblock, dbwiddis, gbbafna, jed326, kotwanikunal, mch2, nknize, owaiskazi19, reta, Rishikesh1159, sachinpkale, saratvemulapalli, shwetathareja, sohami and VachaShah as code owners December 12, 2024 19:49

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Unify precomputation of aggregations behind a common API #16733

Unify precomputation of aggregations behind a common API #16733

msfroh commented Nov 27, 2024 •

edited

Loading

github-actions bot commented Nov 27, 2024

bharath-techie Dec 3, 2024

msfroh Dec 9, 2024

sandeshkr419 commented Dec 3, 2024

github-actions bot commented Dec 12, 2024

codecov bot commented Dec 12, 2024

Unify precomputation of aggregations behind a common API #16733

Are you sure you want to change the base?

Unify precomputation of aggregations behind a common API #16733

Conversation

msfroh commented Nov 27, 2024 • edited Loading

Description

Related Issues

Check List

github-actions bot commented Nov 27, 2024

bharath-techie Dec 3, 2024

Choose a reason for hiding this comment

msfroh Dec 9, 2024

Choose a reason for hiding this comment

sandeshkr419 commented Dec 3, 2024

github-actions bot commented Dec 12, 2024

codecov bot commented Dec 12, 2024

Codecov Report

msfroh commented Nov 27, 2024 •

edited

Loading