Last month I was wrapping up a Java backend project, and the client came back with one more requirement: the product catalog needed "snap a photo, find the same thing."
My first thought was, that has to be Python.
Then I paused. Where did that "has to" come from?
The assumption nobody questions
If you write Java backend code, you probably have this reflex too. The moment an AI requirement shows up, a voice in your head says "this part goes to Python." Image recognition, recommendations, risk scoring: somehow all of it defaults to Python.
The reflex isn't wrong on its face. Model training really does live in Python. PyTorch, PaddlePaddle, and the whole LLM ecosystem are over there, and I'm not going to argue with that. There's nothing to argue about.
But here's the thing: most of the requirements we actually get handed have nothing to do with training.
If you're building image search, what you need is deployment. Throw an image at the system, pull back the closest few from a gallery. That job is about concurrency, latency, stability, and how cleanly it plugs into the system you already have. Which is exactly where the JVM is strong.
We were using the training side's weakness to cancel out the deployment side's strength. That's what made me stop and think.
Two familiar paths, and what each one costs
When a Java project needs AI, there are usually two options on the table.
The first is to stand up a Python microservice and call it over HTTP or gRPC. Looks straightforward, until you count what comes with it: cross-language serialization overhead, two tech stacks, two deployment pipelines, two sets of monitoring and alerting. When something breaks in production, the Java team and the ML team have to debug it together, and just figuring out which side the problem is on can eat half a day. The complexity doesn't go up by one. It gets multiplied.
The second option is to move the business over to Python entirely. That costs even more. You're throwing away the entire stack and infrastructure your team has built up.
I didn't want to accept either one. So I spent a long stretch of time trying a third path: put the whole AI capability inside a single Java process. No cross-language calls, no external service to run.
The result is called ImageSearch4J. It's at version 1.0.0 now, and it's open source.
POST /api/search
│ image (File) + topk + threshold
▼
Main-body detection (PicoDet-LCNet, 640×640)
▼
SANMS refinement (fixes "big box swallows small box")
▼
Per-candidate crop → feature extraction (PP-LCNetV2, 224×224 → 512-dim)
▼
Per-candidate vector search → re-ranking to pick the best subject
▼
Take the best subject's vector, pull back the similar-image list
The whole chain, from image in to results out, is one JAR and one process. No Python. No RPC. No separate vector database to keep alive.
Here's what I want to walk through: the decisions behind it. In my experience the reasoning is worth more than the code, and it's also the part people ask about most.
Three pieces, and they're really the same decision
To make this path work, three problems had to be solved: how to move the model over, what to run inference with, and where to put the vectors for search. That maps to ONNX, DJL, and Lucene.
Moving the model over: choosing ONNX to fully separate training from deployment
Most people treat ONNX as a build artifact. You export it and move on. I think of it as a contract.
It's a Linux Foundation project now. It doesn't belong to PyTorch and it doesn't belong to PaddlePaddle, and neither of them can change it unilaterally. What it defines is a computation graph with no framework attached to it.
What that buys you: the training team can use PyTorch or Paddle, and can switch to whatever framework shows up next year, and not a single line of Java changes. On the other side, the deployment side only needs to understand ONNX. It doesn't care how the model was trained.
The handoff goes from "framework to framework" down to "framework to standard."
Later I ran into a passage on the ONNX site that says almost exactly what I'd been thinking:
We believe there is a need for greater interoperability in the AI tools community. Many people are working on great tools, but developers are often locked in to one framework or ecosystem. ONNX is the first step in enabling more of these tools to work together by allowing them to share models.
Our goal is to make it possible for developers to use the right combinations of tools for their project. We want everyone to be able to take AI from research to reality as quickly as possible without artificial friction from toolchains.
"Without artificial friction from toolchains." That phrase covers it. Cross-language calls, an extra service to run, doubled deployment. None of that is difficulty inherent to the problem. It's difficulty the toolchain created for itself.
A concrete example: the two models ImageSearch4J uses are PicoDet-LCNet and PP-LCNetV2 from the PP-ShiTu family, trained on the Paddle side and exported to ONNX. The entire Java side does inference and nothing else. Swap in a different model, and as long as it can export to ONNX, it drops in.
There's a quieter benefit here too. The ONNX spec requires backward compatibility: a newer runtime has to be able to run an older model. For a project you intend to maintain for years, that matters more than how good it feels today.
Running inference: choosing DJL because this problem was already solved
Someone asked me whether writing your own inference layer in Java is a lot of pain.
It isn't, because you don't have to. AWS has an open-source project called Deep Java Library, DJL. It went public at re:Invent back in 2019. Its positioning is engine-agnostic: one set of Java APIs, and underneath you can plug in ONNX Runtime, PyTorch, TensorFlow, whatever. It's been running on AWS infrastructure for years. DJL Serving is the official model deployment path on SageMaker, with dynamic batching, autoscaling, multi-engine hosting, the whole set.
Put plainly, "running AI inference in Java" is not uncharted territory. I don't need to build an inference engine. I stand on DJL and put my energy into image search specifically.
One security note worth passing along. AWS published an advisory for DJL affecting versions 0.13.0 through 0.36.0, fixed in 0.37.0. This project is on 0.38.0, comfortably past that line.
Where the vectors go: choosing Lucene to delete an entire component
This one took me the longest to settle.
The usual move is a purpose-built vector database: Milvus, Qdrant, or Elasticsearch's KNN. They're all capable and full-featured. But in my situation they share one problem: they add another thing.
Another process, another deployment, another dashboard, another failure point. Expanding the operational surface of the whole system just to get one search feature didn't feel like a good trade.
Apache Lucene is the other road. It's the search core in the Java ecosystem that has survived the longest; Elasticsearch and Solr both build their retrieval on top of it. For what I'm doing, two things about it fit especially well.
First, it's embedded. It's a JAR. Same process as the application, no standalone service, no network hop. That entire operational layer just disappears.
Second, it has had native HNSW vector indexes since 9.0. On top of that I added scalar quantization, squeezing float32 down to int8. The index memory drops to roughly a quarter, and recall stays above 95%. That optimization came essentially for free.
There's a side benefit I only noticed later. A Lucene document can carry a vector field and a text inverted field at the same time. Which means hybrid search (vector plus keyword) already has its foundation in this architecture, without pulling in anything new. I'll come back to that.
Those are the three pieces, and they're doing the same thing: squeeze what has to be coupled down to a minimum, and let everything else stay free. Training decoupled from deployment, model decoupled from engine, AI capability decoupled from infrastructure. That's a line I only saw in hindsight. The three choices look independent, but underneath they follow one rule.
The bug that wouldn't go away: big boxes swallowing small ones
Before I get into this, one note. The algorithm isn't complicated. A few dozen lines. But it's where I spent the most time on the whole project, because figuring out what to fix was far harder than fixing it.
How it showed up
There's a bottle of cola in my gallery. I fed in a photo of it, expecting it to match that bottle.
It matched something else. Specifically, it matched a different image: a close-up of the label.
At first I assumed the model was weak. I tried a few more photos, and the pattern held: whenever the target has a large label, logo, or block of text, the match tends to land on a close-up of that region. Search for a cola bottle, get a label. Search for a drink bottle, get a bottle cap.
That's not a fluke. That's a stable failure mode.
Looking at the actual detection output
The first thing I did was print out what the detector was returning. That explained it immediately.
The detector produced two boxes on that cola photo:
- A large box around the whole bottle, confidence 0.60
- A small box around just the red-and-white label, confidence 0.95
One box is "the bottle," the other is "the label." And since the downstream logic sorts by confidence, the label box at 0.95 wins. The system takes the label and compares it against the gallery, so of course it comes back with label images.
The problem was never "is the search accurate." It was that the wrong thing got picked for searching.
The approaches I tried that didn't work
Once I understood the symptom, I worked through several fixes. None of them held.
Raising the confidence threshold. Push it above 0.6 and the label box is gone. But so is the bottle box at 0.60. You end up with zero candidates. Lower it and the noise gets worse. The road is blocked because the large box is naturally low-scoring and the small box is naturally high-scoring, and there's no single threshold that separates them.
Tuning the NMS IoU threshold. The idea is to merge the two boxes into one. But NMS decides "are these the same object" using IoU, and in a nested case the overlap is tiny relative to the union. You might get an IoU of 0.01. The value is already almost zero, so no threshold I pick can ever touch these two boxes. The criterion itself is pointed at the wrong thing.
Deleting all small boxes and keeping only the largest. I spent the most time on this one. It does keep the bottle, but it throws away the small box's 0.95 along with it. You're left with a 0.60 bottle box that doesn't rank anywhere useful. You've traded one problem for another.
And "keep the largest" doesn't hold up as a rule anyway. The biggest box isn't necessarily the subject. It could be the whole shelf, or the tabletop. In real photos, the background box is often larger than the subject.
The moment it clicked
After every path was blocked, I finally saw where the problem was.
NMS's criterion is overlap. It answers "do these two boxes overlap a lot." But a large box wrapping a small one isn't overlap. It's containment, and that's a different geometric relation.
Containment can have an IoU approaching zero. Overlap always has a large one. Measuring containment with an overlap ruler doesn't work, no matter how you tune it.
So I changed rulers. Instead of looking at overlapping area, compare coordinates directly and ask whether one box falls geometrically inside another.
Once that clicked, the rest was just writing it down. I named it SANMS, short for Structure-Aware NMS.
What it actually does
Three steps:
- Process boxes from largest area to smallest, so the "whole" boxes come out first
- Compare coordinates directly to test whether two boxes are in a geometric containment relation (note: not IoU)
- If a large box fully contains a small one, drop the small one and pass the small one's high score up to the large one
That third step is the part I'm happiest with, and it's what separates this from "just keep the biggest box." It lets the large box pick up two things at once: the scale of the whole, plus the confidence of the part. After the bottle box absorbs the label box, its score goes from 0.60 to 0.95, and it wins the ranking on its own merits. Here you don't have to choose between precision and recall.
The core of it is a few lines:
// process in descending area order
Arrays.sort(sortedIdx, (a, b) -> Float.compare(areas[b], areas[a]));
// large box fully contains small box → absorb, and inherit the small box's high score
if (kbox[0] <= curBox[0] + EPS && kbox[1] <= curBox[1] + EPS &&
kbox[2] >= curBox[2] - EPS && kbox[3] >= curBox[3] - EPS) {
// score inheritance (element-wise max), applied at the moment of absorption
if (scores[curIdx] > scores[kidx]) {
scores[kidx] = scores[curIdx];
}
isContained = true;
}
How it relates to traditional NMS
One thing worth clearing up, because it's easy to misread: SANMS doesn't replace NMS. It layers on top of it.
The detection models in the PP-ShiTu family are end-to-end; the detection head already runs a built-in round of NMS on its output. So I'm not touching that. I add a refinement layer after it.
The two handle different things:
- The built-in NMS handles overlap: two boxes with a high IoU are probably the same object, so dedupe.
- SANMS handles containment: two boxes with a tiny IoU are actually the same object, so merge.
One covers the common case, the other covers the blind spot, and together they converge in layers. In the actual code, a first pass filters candidates by score, then SANMS refines them. It's a loop where the coarse pass protects recall and the fine pass protects precision.
A property I really like
Zero extra parameters. Zero training cost.
It's pure geometric post-processing. No retraining the detector, no new parameter to tune. You can drop it behind the output of any object detection model without worrying about "what value does this dataset need."
That mattered to me, because it means it's a clean algorithmic improvement that adds no tuning burden for whoever uses it.
Why it deserves its own section
Back to that opening scenario. The real problem was never "is the retrieval accurate." It was what the system picks up to search with.
Traditional NMS leaves both the bottle box and the label box in place. The high-scoring small box wins. So the system searches by label, when the user wanted the bottle. That kind of mistake is baked in at step one, and no amount of similarity-algorithm tuning downstream can save it.
And big-box-swallowing-small-box is not a corner case. Product search, fashion search, landmark retrieval: anywhere the target has prominent local texture, this failure mode shows up. Whatever I ran into, anyone building something similar will probably run into too.
SANMS exists to fix one thing: realigning "what was detected" with "what should be searched for."
"No training required" is worth more than a bullet point
Training-free is one of the more important properties of this project, but it's easy to wave off as marketing. Let me actually explain it.
Why can image search be training-free? The root is a distinction that gets blurred a lot: this is a retrieval task, not a classification task.
| Classification | Retrieval (image search) | |
|---|---|---|
| Where the knowledge lives | In the model weights | In the gallery |
| Adding a new class | Retrain; needs data and compute | Add a few images |
| The model's role | A classifier | A general "image → vector" function |
In classification, the knowledge is baked into the weights. To recognize a new class, you retrain. Retrieval is different. The knowledge is in the gallery, not the model. The model only has one job: turn any image into a vector, reliably. A general pretrained model already does this well. It doesn't need adjusting for your data.
That's why training-free holds up.
It leads to a few very practical consequences.
Onboarding a new category means adding images to the gallery. No training, no labeling, no ML engineer required. In this architecture, the "vector update" operation replaces what other systems call "retraining."
The person using it doesn't need to know AI. No loss functions, no learning rates, no batch sizes, no GPU cluster for training. You bring images.
The Java side only ever does inference. Training is heavy work: datasets, distributed runs, hyperparameter tuning, GPU scheduling. Inference is light: load a model, run a forward pass. Because it's training-free, that entire mass of training complexity is deleted up front. This is why the project can stay lightweight. It's the root of it.
It's also where the lineage with PP-ShiTu shows. Same underlying models, same positioning, same training-free property. To move to a different industry scenario, you usually just need a different gallery.
What it looks like today
That's the thinking. Here's the actual thing.
Architecture
┌──────────────────────────────────────────────────────────────────────────┐
│ Spring Boot 3 Application (Single Process) │
│ │
│ POST /api/search │
│ │ image (File) · topk · threshold │
│ ▼ │
│ ┌─────────────────────── ImageSearchService ───────────────────────┐ │
│ │ @Async("aiInferExecutor") │ │
│ │ │ │
│ │ ① PipelineService.process() — locate the "best subject" │ │
│ │ ├─ MainBodyDetectionService.predict() │ │
│ │ │ PicoDet-LCNet detection → Top5 pre-filter → threshold │ │
│ │ │ → SANMS: big box absorbs small + score inherit │ │
│ │ └─ Per candidate: crop → PP-LCNetV2 embedding → vector search │ │
│ │ tie-break: larger area wins → set as match │ │
│ │ │ │
│ │ ② Match vector → retrieve full similar list (similarList) │ │
│ └──────────────────────────────────────────────────────────────────┘ │
│ │
│ Inference: AWS DJL 0.38 ── ONNX Runtime 1.29 ── PicoDet / PP-LCNetV2 │
│ Retrieval: Apache Lucene 9.12 ── HNSW + scalar quantization │
└──────────────────────────────────────────────────────────────────────────┘
Pipeline and API
There's one main endpoint:
POST /api/search
Content-Type: multipart/form-data
| Parameter | Type | Required | Default | Notes |
|---|---|---|---|---|
image |
File | yes | — | The image to search with |
topk |
int | no | 10 | Number of results; falls back to 10 if below 1 |
threshold |
float | no | 0.4 | Similarity threshold for vector search; falls back if out of range |
A curl example:
curl -X POST http://localhost:8080/api/search \
-F "image=@./cola.jpg" \
-F "topk=10" \
-F "threshold=0.4"
The response looks like this (real capture):
{
"state": 1,
"code": 200,
"message": "Success",
"data": {
"match": {
"name": "康师傅冰红茶",
"path": "142/2.jpg",
"md5": "96b9f53217b116af6b767ba033e74b9f",
"score": 0.8343468
},
"similarList": [ ... ],
"candis": [ ... ],
"rect": [499.86, 9.44, 801.02, 945.43]
}
}
Two fields are worth explaining. match is the best match the system settled on, and similarList is the full similarity list, ranked. rect is the bounding box of the best-matching subject, and candis is the other candidate boxes besides it, each with its detection confidence.
Why split it into two groups? Because match answers "what is most likely in this image," and similarList answers "what in the gallery looks similar." Different meanings, so the front end can present them separately. And keeping rect apart from candis lets a user see directly on the image why the system decided what it did. Draw the accepted subject and the rejected candidates in two colors and it's obvious at a glance.
One detail worth mentioning: even if nothing in the vector store matches, rect and candis still come back populated. That means the front end can say "a target was detected, nothing matched in the gallery" instead of showing a blank page.
The interface
The project ships two zero-framework static pages. No front-end framework at all, plain HTML and vanilla JS.
The search page accepts images three ways: local upload, image URL, or pasted screenshot. All three are converted to a File object before submission, so the backend only ever sees one image field. No adaptation needed on the API side. When results come back, the preview image gets annotated automatically with the best-match box (green, solid) and candidate boxes (orange, dashed), with a match card and similar-image list below.
The other page is for vector updates. One thing to be straight about: it's a demo page right now. It exists to show the interaction shape for incremental vector updates: bulk upload, stats, logs. Making it a production vector management console would mean filling in its backend endpoints. The docs say so too. I have no interest in dressing it up as a "full admin panel."
Some things that don't belong in a README
These are the parts I cut from the README but still think are worth saying out loud.
It sits where Spring AI sits
I kept looking for an analogy to explain what this project actually is. What I landed on: ImageSearch4J is to image retrieval in Java what Spring AI is to LLMs in Java.
Let me be precise about it. Spring AI doesn't provide models. It wires LLM capability into the Spring ecosystem so developers call it the way they're used to. This project is the same shape. The models aren't mine, the inference engine isn't mine, the search engine isn't mine. What I did was connect them, and connect them into a form a Java developer can use without friction.
The difference is granularity. Spring AI abstracts "how to call an LLM," across multiple vendors. I abstract "how an entire image-retrieval pipeline runs," for application developers. Same layer, same kind of job, not the same shape.
Honestly, I kept this analogy out of the README. For a project that just went open source, opening with "it's basically the X of Y" reads as inflating yourself. But in a setting like this, it's a useful coordinate.
The positioning is a lightweight base, and it isn't changing
This project is positioned as a lightweight base, and I intend to keep it there.
The reason is simple: something huge scares people off. A newcomer clicks in, sees dozens of modules and a screenful of config options, and closes the tab. But the whole value of a base is that people are willing to use it and can get it working.
So I gave myself a test. It's crude but it works: would this change add steps between "clone" and "first search result"? If yes, it doesn't go into the main line, however reasonable it is. It gets written down instead.
More features and a lower barrier often pull against each other. I try to pick the latter. If this project ever genuinely needs to grow (large galleries, clustering, multi-tenancy), that becomes a new project. The old one shouldn't carry the new one's weight, and the new one shouldn't inherit the old one's restraint. That's better for both.
Some things are "not done" on purpose
A few examples.
The main-body detection threshold is hardcoded. It isn't exposed as config. It's an internal algorithm parameter, and handing it to a secondary developer to worry about is the wrong place for it. Externally there's one threshold, and that's enough to control retrieval precision. The more parameters you expose, the more decisions the user has to make, and the higher the barrier.
Hybrid search is the same story. As I mentioned, Lucene makes vector-plus-keyword hybrid search nearly free here. Write name into a text field, combine a KNN subquery with a term subquery using a BooleanQuery, and the fusion happens inside a single query. But I didn't put it on the default path. A new user shouldn't have to understand "how two retrieval modes fuse" before they've gotten their first result back. The capability stays in the architecture. Whether to turn it on is left to someone experienced.
Both fall into the same category: possible, but off by default. Not a missing capability. Complexity deliberately parked where it belongs.
What it doesn't do
A project that only lists what it can do isn't very trustworthy. So, what it can't do right now.
It's not for training. As covered above, it only handles inference and retrieval. If you want to train or fine-tune a model, that's back on the Python side.
It's a different tradeoff from Elasticsearch. If your company already has an ES cluster, then ES 8.x's built-in vector support plus your existing infrastructure is probably the easier call. This project takes the embedded route, and it wins on being standalone, light, and needing no extra ops. But those two roads trade off differently. Neither is simply better.
Hybrid search is off by default. If you want it, wire it up yourself.
Saying this up front beats letting someone find out the hard way. A project with clear boundaries is easier to trust.
Quick start: getting it running locally
Everything above was "why." This section is "how." Following these steps, you go from nothing to your first search result in about ten minutes.
Requirements
- Zulu JDK 17 or later
- Maven 3.8 or later
That's it. No Python, no separate database to install.
Step 1: Prepare a gallery
The project doesn't bundle gallery data. You supply your own. The official PP-ShiTu sample dataset drink_dataset_v2.0 is a good fit. It's full of beverage photos, which makes for a clean demo.
# download
wget https://paddle-imagenet-models-name.bj.bcebos.com/dygraph/rec/data/drink_dataset_v2.0.tar
# extract
tar -xvf drink_dataset_v2.0.tar
# rename to image_gallery
mv drink_dataset_v2.0 image_gallery
# move to your home directory
mv image_gallery ~/
The gallery needs to live under your home directory, at ~/image_gallery/. Once it's there, the structure looks like this:
~/image_gallery/
├── gallery/ # the gallery itself
│ ├── 142/ # subdirectories by category id
│ ├── 164/
│ ├── ...
│ └── drink_label_all.txt # label file
├── test_images/ # official test images, handy for a first search
└── vector_index/ # index dir, generated on first start; absent initially
That drink_label_all.txt is the label list. Each line is one image mapped to a category name, tab-separated. When building the index, the program reads it to know what each image is called.
What if you put it in the wrong place? You won't be left guessing. On startup the app checks for that directory. If the gallery is missing, it prints the download URL to the log and exits on its own, rather than dumping a stack trace for you to decipher.
Step 2: Start it
git clone https://github.com/TommysLee/ImageSearch4J
cd ImageSearch4J
mvn spring-boot:run
First startup does two extra things, both automatic.
One, it releases the models. The two ONNX models are bundled in the project under resources/models/:
-
picodet_lcnet_x2_5_640_mainbody.onnx(main-body detection, ~28 MB) -
general_PPLCNetV2.onnx(feature extraction, ~18 MB)
On first start they're copied to ~/models/. Later starts use them directly, no repeated copying. You never download a model file by hand.
Two, it builds the index. If ~/image_gallery/vector_index/ is empty, the program scans the gallery, extracts features image by image, and writes the vectors into the index. This step is CPU-heavy and takes a while on a large gallery. There's a progress indicator in the console. When it finishes, it's done.
One more thing: once the index exists, new images are picked up through NRT (near-real-time) refresh: visibility refreshed every second, commit every five minutes. So an image added through the API becomes searchable without a restart.
Step 3: Open a browser
The default port is 80:
http://localhost/index.html
On Linux or macOS, port 80 is privileged and a normal user can't bind it. In that case, start on a different port:
mvn spring-boot:run -Dspring-boot.run.arguments=--server.port=8080
Then go to http://localhost:8080/index.html.
Upload a photo of a drink and you'll see the detection boxes and the match results. If you don't have a suitable image handy, ~/image_gallery/test_images/ has the official ones. Pick any.
Step 4: Call the API
Beyond the UI, you can call the endpoint directly. Assuming port 8080:
curl -X POST http://localhost:8080/api/search \
-F "image=@~/image_gallery/test_images/xxx.jpg" \
-F "topk=10" \
-F "threshold=0.4"
What comes back is the structure from the section above. If you're writing your own front end, parse against that.
About onnx-engine-path: the config that trips people up
There's a setting called ty.onnx-engine-path, defaulting to:
ty:
onnx-engine-path: "${user.home}/.djl.ai/onnx/win-x64"
First, the conclusion: this parameter is not required. The program runs fine without it. But it's worth setting in production. Here's what it's for.
Every time ONNX Runtime loads, it extracts its bundled native libraries. On Windows that's three files: onnxruntime.dll, onnxruntime4j_jni.dll, and onnxruntime_providers_shared.dll. By default it extracts them to a temp directory and cleans up when the process exits. The problem is in that cleanup. It goes through File.deleteOnExit, and that method can only delete empty directories. So every run leaves a bit more behind in the temp directory, and over time you accumulate junk on disk.
Pointing at a fixed native-library directory changes this: extract there once, then use it directly on every later run. No repeated extraction, no leftover junk. That's why the parameter exists.
Now the part that's genuinely confusing. The first time you see this config line, you'll probably wonder: where does the stuff in that directory come from? I never created it.
The answer takes a small detour: it comes from onnxruntime.jar.
Specifically, the com.microsoft.onnxruntime:onnxruntime dependency from Maven (version 1.29.0 in this project). Inside that jar, native libraries for every platform are packaged by platform, at:
ai/onnxruntime/native/<platform>/
The directories and files by platform:
| Platform | Directory in jar | Files |
|---|---|---|
| Windows x64 | ai/onnxruntime/native/win-x64/ |
onnxruntime.dll, onnxruntime4j_jni.dll, onnxruntime_providers_shared.dll
|
| Linux x64 | ai/onnxruntime/native/linux-x64/ |
libonnxruntime.so, libonnxruntime4j_jni.so
|
| Linux aarch64 | ai/onnxruntime/native/linux-aarch64/ |
libonnxruntime.so, libonnxruntime4j_jni.so
|
| macOS (Apple Silicon) | ai/onnxruntime/native/osx-aarch64/ |
libonnxruntime.dylib, libonnxruntime4j_jni.dylib
|
So if you want to set this parameter, you have to extract the native libraries for your platform out of the jar yourself, put them where you want, and point the config there.
The jar is in your local Maven repository at ~/.m2/repository/com/microsoft/onnxruntime/onnxruntime/1.29.0/. For Linux x64, one command does it:
unzip -j ~/.m2/repository/com/microsoft/onnxruntime/onnxruntime/1.29.0/onnxruntime-1.29.0.jar \
"ai/onnxruntime/native/linux-x64/*" \
-d "$HOME/.djl.ai/onnx/linux-x64"
On Windows, swap linux-x64 for win-x64, and so on. Then set the config to match:
ty:
onnx-engine-path: "${user.home}/.djl.ai/onnx/linux-x64"
Why does the default say win-x64? Because my development machine runs Windows and I set it for my own box. It's not a universal value. When you deploy on another platform, change it to the matching directory. And if you don't, that's fine too. The program falls back to the default behavior and runs anyway, just through the temp-directory path.
Other things worth knowing
A few more points that may help during setup.
The config file is src/main/resources/application.yml. Port, model directory, index directory, thread pool sizes, all in there with fairly plain names.
To use your own gallery, just prepare a directory with the same shape: gallery/ holds the images (subdirectories are fine), gallery/drink_label_all.txt holds labels, one line per image as relative/path<Tab>category. After switching galleries, delete the old vector_index/ so it rebuilds.
Images show up in the page directly because a file:${user.home}/image_gallery/ mapping is registered for static resources. So the gallery directory doubles as the image access root.
Where this lands
What I most wanted out of this project was to add one more option to the list.
A Java developer hitting an AI requirement doesn't have to switch to Python, and doesn't have to stand up a Python microservice. That path works now: the model travels as neutral ONNX, inference goes to DJL, retrieval goes to Lucene, and Spring Boot assembles it. The point is that each part plays to its strength, and none of them has to accommodate the others.
I've long felt the ideal state for AI is one where application developers never notice it's there. The model loading, the object pool, the index refresh, the concurrency: all of it goes inside the Service layer, and business code only ever faces one clean method call:
@Service
@RequiredArgsConstructor
public class ProductSearchService {
private final ImageSearchService imageSearchService;
public SearchResult search(byte[] imageBytes) throws Exception {
BufferedImage image = ImageIO.read(new ByteArrayInputStream(imageBytes));
return imageSearchService.search(image, 10, 0.4f).get();
}
}
(Note that search is annotated @Async("aiInferExecutor") and returns a CompletableFuture, so remember .get() or .join() to retrieve the result.)
Put the AI in the foundation. Leave the business to the people writing it. That's what this project is trying to do.
It's open source under Apache License 2.0. Commercial use and secondary development are both fine. If you've also been bothered by the missing AI piece in the Java ecosystem, come take a look, and issues and PRs are welcome. If it's useful, a star is nice. If not, that's fine too. With something like this, every person who learns Java can do it too is a little more value.
Repository: https://github.com/TommysLee/ImageSearch4J




Top comments (0)