Photo In, Form Out: How Multimodal AI Quietly Solved the Worst Part of Selling Things Online
There is a cordless drill on a shelf in my garage that I have been meaning to sell since about last spring. It works fine. The battery is missing, which I would have to disclose, and the model number on the sticker has worn down to roughly four legible characters. To the right person, it is worth maybe forty dollars. I have not listed it, and at this rate I am not going to, because listing it means twenty minutes of typing that I resent more than I want the forty dollars.

That resentment is one of the more underrated problems in consumer tech. It has finally started to move.
The photo was never the hard part.
Ask people why they haven’t sold the stuff piled in the spare room and most of them will say something about photos. Bad lighting, no clean wall, the shot always comes out crooked. It sounds right. It is also mostly wrong, and has been since phone cameras got good.
The real friction sits underneath the photo. It is the form.
Open a new listing on eBay, and you get a title field with a character limit, a category tree you have to guess your way down, and then a stack of item specifics: brand, model, color, material, size, condition, sometimes a dozen more depending on where in the tree you landed. Etsy wants tags and attributes. Poshmark wants size and brand. None of it is difficult, exactly. It is fiddly and repetitive and dull, and it has to be done once per item, forever.
It seems like no harm to skip those fields. It isn’t. Without images, the person searching for a specific cordless drill won’t be able to find a listing with “drill” in the title, and even if the category were set to the default, there would be no images to show. This could be the perfect picture. No one is paying attention to it.
But it’s the same with clothes: the majority of casual sellers begin with clothes. The one I put up, a wool coat, no brand, no size, and no material, a comment on the sizeable stain on the hem that you had the integrity to include, is a good listing because it won’t be found by the person typing a brand and a size into the search bar. If the fields above the description are empty, then it is not helpful to be honest about your description.
The majority of second-hand selling tips are upside down. Sourcing does not necessarily limit it; it is not a constraint. It is not a problem of the camera. Data Entry is, and no one writes blog posts about data entry.
Why this turned into a multimodal problem
There’s a reason for this: image recognition alone was never the solution to this, and it matters. You’ve been told what you know when a classifier analyzes a photo and provides you with 94% confidence that it’s a power drill. Has not completed one item.
There were two changes, but both are recent.
The first is that models are now taught to analyze a series of photos as a body of evidence as opposed to one image at a time. The brand is on the box in the third shot. The part number is stamped on the bottom and only appears in the fifth. The scuff on the housing is only found on the second and is not on any other housing. It takes about ten seconds for someone to pick up the thing and turn it over, and then write one coherent description from that. The photograph reading together is much closer to that than reading the pictures one by one.
The second is that there is some output restriction. It’s easy to create a nice paragraph about a drill, and little is useful. What a listing requires is a category that is in that market’s tree and the attribute values the platform will recognize. Unconstrained generation makes up attributes that sound plausible, and thus it gets the listing bounced. Having the real field list in the platform and only what’s allowed to be inputted is boring engineering and is what makes any of this usable.
What that does to the arithmetic
This is when it gets fun for those wanting to sell items. If it takes 20 minutes to post, then anything under 30 dollars isn’t worth posting; and you can see that’s a line of arithmetic that explains a lot of the usable that’s in bins in hallways and not going to somebody who wants it. Listing AI tools have pushed the per-item cost down to something closer to a minute, which quietly puts the cheap end of the pile back on the table. That is a bigger shift than the demo videos make it look.
The process from tool to tool is generally the same. Take some photos of the object, then upload the photos to the set, wait for a minute and read what it says, fix it if something is wrong, and publish it. Most of them also edit the pictures, making the background white or transparent and fixing pictures taken at an angled level.
This is one difference that is not given enough focus. The background is cropped out. It doesn’t regenerate, and its object is never redrawn. When the quality of the product is paramount, as in the case of used goods, where the buyer pays and is purchasing the actual product with all its actual imperfections, then those tools that make the product look good without making a show of it are selling something that won’t ever get to them.
Where it still stops short
Some honesty about the limits, since the gap between the pitch and the reality is usually where people get annoyed.
- Publishing is mostly still copy and paste. I have not used anything else like Etsy. Once you connect to a shop, your listing will be published instantly, complete with pictures. With eBay, Poshmark, Mercari, Craigslist, and Facebook Marketplace, you create your listing and paste it yourself. Not painful, but not the one-click flow that some categories suggest.
- A suggested price is a starting point, nothing more. These are not “re-pricing” engines. When you create a range, you immediately have a range, and it’s yours to do with as you please.
- You still have to read it. An ever-so-worn model number can be read with absolute certainty. So can a situation call on a defect the camera leveled out. If you don’t want to be embarrassed to have it wrong in a disagreement, check with your own eyes.
None of that is a dealbreaker. It just means the software does the typing and you keep the judgment, which is roughly the right division of labor anyway.
The bottom line
The attention-grabbing AI stories are the ones that are about “new things. The stories that get attention are the ones that are about “new things” with the AI. Video from a still image, voice from a text file, talkative face from a still image. The other story that might be more useful this year would be the opposite: boring structured work managed well enough so that the little stuff is worth doing all over again.
I have a drill in my garage, but it’s not in use. The reason why it is there is dwindling.