This patent's title says LiDAR but its classification says neural network. UATC's grant US11221413B2 ("Three-dimensional object detection," issued January 11, 2022; inventors Ming Liang, Bin Yang, Shenlong Wang, Wei-Chiu Ma, Raquel Urtasun) carries G06N 3/02, 3/04, and 3/08 — network architecture, layered structure, and training — right alongside the G01S LiDAR codes. The LiDAR is the input; the claimed invention is the network that turns points into boxes.

Claim 1 is not a generic "detect objects with a neural net" claim — it recites a specific four-system fusion architecture, and the distinctive ingredient is a map. The claim names a camera system, a LiDAR system, and a "map system configured to provide geographic prior data…determined by a machine-learned map estimation model" that takes LiDAR as input and outputs "geometric ground prior data and semantic road prior data." Then a "fusion system" does three things: it "modifies, based on the geographic prior data, the LIDAR point cloud data…into map-modified LIDAR data having a plurality of layers, each layer representing a two-dimensional view"; it fuses "image features from the image data with LIDAR features from the map-modified LIDAR data"; and it generates "a feature map comprising the fused image features and the LIDAR features." A detector then reads 3D objects off that feature map. The inventive content is using learned map priors to reshape the LiDAR before fusing it with camera — not detection alone.

The dependent claims make the map-modification mechanism concrete, and it is genuinely clever. Claim 2 modifies a "bird's eye view representation" of the LiDAR "by subtracting ground information…from the LIDAR point cloud data" — flattening out the ground so the network sees objects, not terrain. Claim 7 defines the ground prior as "a point-wise representation of ground height," and claim 8 spells out the operation: "replacing a value along the Z-axis for the one or more initial points with a value along the Z-axis from the point-wise representation of ground height" — literally re-zeroing each point's height against the learned ground, so a curb and a car are measured relative to the road surface rather than the sensor. Claims 3, 4, and 11 add the semantic side: "extracting a semantic road region mask from a high definition map and rasterizing the semantic road region mask onto the LIDAR point cloud data as a binary road mask channel" — a drivable-region channel concatenated onto the data so the detector knows where road is. Ground subtraction plus a road-mask channel is the recited map-priors-into-LiDAR pipeline.

“Generally, the disclosed systems and methods implement improved detection of objects in three-dimensional (3D) space.”— U.S. Patent No. 11,221,413 source

The camera-LiDAR fusion itself is a recited mechanism, not a hand-wave. Claim 6 requires "one or more fusion layers" that "fuse image features from the image data at a first level of resolution with LIDAR features…at a second level of resolution that is different" — multi-scale, cross-modal fusion. Claim 19 names the operator: "one or more multi-layer perceptrons" whose first portion extracts "source data points associated with the map-modified LIDAR data given a target data point associated with the image data" and whose second portion encodes "an offset between" them. That offset-encoding MLP is the continuous-fusion operator the abstract references — it associates a camera pixel with nearby LiDAR points by learning their geometric offset, which is how image and point-cloud features get aligned despite living in different spaces. Claim 5 adds robustness: when the map is unavailable, a "map estimation system" generates "estimated geographic prior data," so the pipeline degrades gracefully off-map.

Claims 13–15 close the autonomy loop — the detector outputs "a bounding shape representative of a size, a location, and an orientation of each" object (claim 14), a "motion planning system" consumes the detections, and "a vehicle control system" executes the plan (claim 15). CPC-wise, G01S 17/89 (LiDAR imaging) and 17/931 (LiDAR for vehicles) mark the input, G06N 3/02/04/08 mark the learned core, and the result ties to vehicle position control. The independent method (claim 16) and autonomous-vehicle (claim 20) forms re-fence the same map-modified-fusion pipeline.

For the control beat, what's striking is that this is one node in a deliberate multi-grant family: the same title and inventors recur across US11500099B2 (2022), US11768292B2 (2023), and US12051001B2 (2024). That cadence is a strategy, not an accident — a chain of continuations re-fencing 3D detection as the art moves, keeping a live, evolving claim set on the same core capability across four years. The technical through-line — learned map priors fused with camera and LiDAR — is exactly the Urtasun-group lineage that produced foundational deep-LiDAR-detection work.

From a portfolio angle, that continuation chain is the real asset, not any single grant. A reader auditing UATC's (and successor Aurora's) perception moat should read the family as a unit: each child can carry slightly different scope, and the latest live continuation defines what is actually enforceable today. A quiet continuation is a loud signal — here, of sustained investment in map-aware learned LiDAR detection.

Caveats. Deep 3D detection from LiDAR is among the most-published autonomy topics; any single claim faces heavy prior art, and granted scope hinges on the specific steps in claim 1 — here, the map-prior modification (ground subtraction, road-mask channel) and the multi-resolution MLP fusion. An accused system likely has to use learned map priors to modify the LiDAR before fusion to read on claim 1. The family's value is that it spreads the bet across formulations. Read the latest continuation's independent claim with claims 8 and 19 for current scope.

For the file: a foundational, multi-generation learned-LiDAR-detection family whose claim-1 hook is map-priors-into-LiDAR fusion, not detection alone. Pull US11221413B2 and its 2022–2024 continuations in the patent record, read the newest child's claim 1 with claims 8 and 19, and treat the chain — not the first grant — as the enforceable fence.