Not a single word about when/if they'll actually release the weights for this, or am I missing it somewhere?
user43928•Jul 21, 2026
Considering Qwen-Image-2.0 weights have not been released either, it unfortunately looks unlikely.
vachina•Jul 21, 2026
Why release it just so Cursor/Azure/Amazon can profit off of it? Unless OpenAI actually opens it up fat chance.
vunderba•Jul 21, 2026
After Z-Image and the original Qwen-Image - I think they've pivoted completely to closed source. From the images I've seen, Qwen-Image 3.0 is just a subpar equivalent to other proprietary models like gpt-image-2 and nb-pro.
Mashimo•Jul 21, 2026
> Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.
Impressive.
Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?
woadwarrior01•Jul 21, 2026
Krea-2-Turbo. I've even got it working locally on my M5 iPad Pro.
Mashimo•Jul 21, 2026
thanks mate. Sadly not supported yet by Invoke, but I will take a look.
Izmaki•Jul 21, 2026
> We implemented safety measures across the full model development lifecycle.
Any suggestions for the best open, non-opinionated model?
woadwarrior01•Jul 21, 2026
It is a reasonably non-opinionated model. My usual test is to ask these models to generate comic book and cartoon characters that hosted image generation models refuse to generate. I think that text is just CYA legalese.
Izmaki•Jul 21, 2026
What's your favourite (online or offline) model for image generation?
bitexploder•Jul 21, 2026
You can also ablate it.
spwa4•Jul 21, 2026
Appears to be closed-weights entirely. No word at all on any weights release.
weird-eye-issue•Jul 21, 2026
The meta keywords in the HTML is very interesting. 100+ references to NSFW topics such as hentai, nudes, etc.
walrus01•Jul 21, 2026
Wow you really weren't kidding. Fully automated AI slop tentacles for everyone!
spiderman pointing at spiderman meme trans wallpaper xxx thicc hentai
WithinReason•Jul 21, 2026
probably due to misspellings as qwen stefani
alex_suzuki•Jul 21, 2026
> ben 10 hentai game where qwen gets drunk
erm... what?
walrus01•Jul 21, 2026
> qwen ten hentai, qwen ten xxx, qwen tennason footjob, qwen tennison hentai, qwen tennyson, qwen tennyson ass expansion, qwen tennyson giantess pussy, qwen tennyson hentai, qwen tennyson porn, qwen tft, qwen the milk maid, qwen the milkmaid, qwen the parts over, qwen the wolf witcher, qwen tire, qwen tneyson porn, qwen to go feelies youtube, qwen tokenizer java,
bbor•Jul 21, 2026
Whelp today’s a unique day on hacker news, wow! Didn’t expect to read that at my desk lol
archon1410•Jul 21, 2026
Surely it was meant to be "Gwen" (the protagonist's cousin and the other main character in the show). I guess it was somehow (incorrectly) assumed to be a misspelling of Qwen and included in the tags. Or perhaps many people were misspelling Gwen as "Qwen" and it was all hoovered up.
pdpi•Jul 21, 2026
Keyword typosquatting I guess?
goobatrooba•Jul 21, 2026
Wow, crazy things - "shemale porn", "spider fucks venom", and various porn URLs like tushy.com. What the hell did they do to their AI? Is that the training corpus or intended use case?
walrus01•Jul 21, 2026
> Is that the training corpus or intended use case?
[whynotboth.gif]
chmod775•Jul 21, 2026
Bonus points for also covering misspellings like "pregbant".
Is there such a thing as Reddit "unslop"? Thought it went without saying.
hedora•Jul 21, 2026
Please don’t violate community guidelines like this. All HN threads look like this. Don’t make me report you to a submod.
> Please don't post comments saying that HN is turning into Reddit. It's a semi-noob illusion, as old as the hills.
sgc•Jul 21, 2026
> Don’t make me report you to a submod.
This is pathologically outside scope. I don't think I have ever seen somebody threaten somebody else on hn before. The 'don't make me do it to you' abuser trope is next level.
user-•Jul 21, 2026
dont make me report you to a sub sub mod!
user_7832•Jul 21, 2026
That's racist!
(See kids, it is possible to fight memes/racism with memes! And well... yeah, this really is racist.)
lukan•Jul 21, 2026
That is hilarious as porn is illegal in China. But I guess pron SEO is allright, if it is against the west.
weird-eye-issue•Jul 21, 2026
Western search engines simply ignore the meta keywords tag for over a decade now
lukan•Jul 21, 2026
So what is the point then?
weird-eye-issue•Jul 21, 2026
There is no point. It's from people that have no clue which is most people that do search engine optimization
That said it's possible search engines in China or other countries might use it, but it's very easy to game so it doesn't really make sense
vitalyan8184•Jul 21, 2026
uh, are you implying that porn is bad somehow? I can show you a hundred articles that say only nazi incel chuds believe so.
That's hilarious. These keywords are applied globally, even on pages like https://qwen.ai/usagepolicy, where they very kindly ask you not to use their products for sexual content.
who made qwen stefani's dress
spiderman into the spider verse did qwen meet Peter
qwen stacey porn
do blake shelton and qwen stafani have children together
Which LLM do you think they used to generate the keywords
postalcoder•Jul 21, 2026
I've been thinking about this too much because it's so ridiculous and funny that I had to at least try wrapping my head around it. My best guess is this:
If you look at past snapshots at archive.org, you notice that the meta keywords are growing like an append-only list, which means it's probably part of some messed up seo pipeline. The other clue is that it's stuffing the meta keywords which apparently only Yandex uses as a search signal[0].
The list has 3882 entries. A lot of them are clustered and look like auto-complete results. But a bunch of them look like hyper-specific, misspelled search results (e.g. "145 gwen rd cheshire ct"). Google's webmaster tools doesn't provide distinct queries like that, but yandex's does[1].
My best guess of what's happening is that Qwen is monitoring it's search queries in Yandex, dumping that list into a serp service that scrapes yandex's autocomplete suggestions, and then taking that list and dumping it into their meta keywords.
It explains the urls in the list (ppls using search engines like address bars), the seeming fixation around certain topics (which usually starts with a misspelling), and the random one-off queries.
Why would they added URLs for porn websites to their keywords? Like literally the names of websites with ".com" and such at the end.
WhereIsTheTruth•Jul 21, 2026
Brainrot is the new opium war
cyanydeez•Jul 21, 2026
how the turn tables
user43928•Jul 21, 2026
I assume they use a SEO tool that automatically adds these meta keywords to optimize for some search engines.
It apparently adds common search terms that contain words like "qwen". This evidently includes possibly mistyped searches for "gwen" or "ben" in a NSFW context.
Maybe someone knows more about how such SEO tools work, and where they pull the data from.
spiderfarmer•Jul 21, 2026
Sounds like “keyword shitter” (genuine tool)
01284a7e•Jul 21, 2026
What value does the 'keywords' meta tag even have these days? 77+ KB of crap added to the page weight. Web development is full of idiots.
j0ej0ej0e•Jul 21, 2026
Maybe it's completely intentional and it's making a comeback at least with LLMs? Google say it has no value, but LLMs probably genuinely use it.
drakythe•Jul 21, 2026
Interesting too that they would include "ai friend" when China just added restrictions on AI "partners" such that many services stopped offering them rather than try to adhere to the restrictions.
ssalka•Jul 21, 2026
I think the Qwen team is probably aware that the NSFW community is very quick to adopt any new image gen model (see: Civitai). So, it seems like a good SEO approach to try and surface their model in search results that said community is likely already checking.
rvz•Jul 21, 2026
Midjourney already knew that image generation was going to zero. Again yet another reason why the model was never a moat in the first place.
amelius•Jul 21, 2026
The moat is the training data. But somehow we've collectively decided that it's not.
treetalker•Jul 21, 2026
The red-dress woman's vestigial pinkie toes …
postalcoder•Jul 21, 2026
They must have trained on GPT Image 1 outputs. The yellow tint is unmistakable.
GPT Image 1 ended up with a yellow tint without training on another image-generation model's output. It's just that humans like pictures with a soft sunset glow, and this is a very easy global signal for a preference model to pick up on, and for a image-generation model to imitate. So optimizing for aesthetic appeal makes everything slightly tinted by default, unless you make sure to countersteer.
postalcoder•Jul 21, 2026
Fair point!
Grimblewald•Jul 21, 2026
Didnt the piss tint come after their big studio ghibli heist?
dannyw•Jul 21, 2026
AI-generated images are part of the web now, if you're doing ordinary web scraping, you can't avoid training on generated images.
hugmynutus•Jul 21, 2026
yellow/red tint is an extremely common problem not matter the photograph source you train on
source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle
Wow, it displays Korean properly without breaking.
But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
sheept•Jul 21, 2026
I find it mildly interesting how your comment repeats itself but with different phrasing
jdw64•Jul 21, 2026
It's because of the structure of Korean. I think in Korean first, so I end up translating it directly.
When I want to emphasize something, I tend to repeat it
mahimai•Jul 21, 2026
interesting
mynti•Jul 21, 2026
To me it feels so weird that people are trying to push these model for online shopping like "here is how this dress/shirt/pants would look on you". But these models will always make the clothes fit your body and show you in flattering light and so on. How the actual garment fits is still as elusive as before these tools
walrus01•Jul 21, 2026
The cynic in me says this is just an advancement and logical continuation from the known problem of sycophantic behavior in text-to-text chatbot format LLM, to image generation models.
inerte•Jul 21, 2026
To me the tools are newer, but fit people in a controlled context to sell clothing is as old as photography itself.
number6•Jul 21, 2026
Sounds exactly like something the marketing department would use to up the convertion rate
viraptor•Jul 21, 2026
There's definitely a predatory "everything will look good", but you can also leverage exactly that part for your own purposes - I've done pre-shopping a couple of times by asking for something like "a grid of 9 versions of my photo, wearing different types of X that look good". Definitely helped with choosing a good style.
khurs•Jul 21, 2026
As long as the Product Manger gets a bonus/promotion then it's a success.
TylerE•Jul 21, 2026
To be fair in many cases the actual clothes aren't much better. I've had two pair of the same pants, same brand, same size fit noticeably differently.
morningsam•Jul 21, 2026
Maybe this will lead to more business for tailors doing alterations, assuming the clothes people end up buying are expensive enough to justify it.
jimnotgym•Jul 21, 2026
I understand they are already doing decent business since the advent of the magic weight loss pen
epolanski•Jul 21, 2026
I have a similar use case at work for previewing construction material and such in our catalogue applied to user uploaded images.
Results are mixed, expensive, but it really feels you're few months off the next improvement to really nail it. It's already good enough.
Wonder what Qwen image will provide over nano banana.
vunderba•Jul 21, 2026
Well from the images I've seen they seem to be pushing text-heavy "infograph" capabilities pretty hard. Outside of that, I have serious doubts that it'll be better than other proprietary models like NB Pro and gpt-image-2.
k2enemy•Jul 21, 2026
I've noticed something similar in Facebook marketplace ads for used furniture. Most of the images are AI generated to look like a Pottery Barn catalog, then the last image will be the actual item, full of scratches and other damage, sitting in a messy garage.
aembleton•Jul 21, 2026
Estate agents are also doing it; making interiors of houses look very different to how they really are.
jeremyjh•Jul 21, 2026
Yes and even if all they do is ask it to stage an empty room picture with furniture and decorations it will make improvements like adding a doorway to a room that doesn’t exist.
agile-gift0262•Jul 21, 2026
The worst I've seen changed the view out a window from an alley to a private garden. Most also add natural light that isn't possible, and in some cases I'm convinced the generate image is depicting the space as larger than it actually is, with furniture and spaces that wouldn't fit in reality
hiccuphippo•Jul 21, 2026
I once heard about a company that specialized in making furniture about 10% smaller than normal for use in show rooms to make the spaces look larger. Not surprised they do it with AI if they can.
mikepurvis•Jul 21, 2026
Wide angle lenses have been a thing in real estate photography since forever, but you would always have the reference of the furniture to ground your perception. Having furniture be slightly downsized is diabolical.
masfuerte•Jul 21, 2026
This is a thing in show homes on new-build estates in the UK. British houses are often tiny.
skinfaxi•Jul 21, 2026
Seems like false advertising.
spaceman_2020•Jul 21, 2026
This should really be illegal
QuantumGood•Jul 21, 2026
And knocking out the messy garage background to replace it with a showroom is the easiest prompt.
mikepurvis•Jul 21, 2026
I've started to see this on Etsy and Wayfair too, where there will be a listing that is clearly just MDF flatpack being resold from China, but the AI-generated images wildly exaggerate the proportions of it.
Ironically, ChatGPT is decently good at ferreting these out. Like I sent it a screenshot of that listing and it not only helped me find where the original item was for sale, but also pointed out how the dimensioned diagram shows it as being just 49" tall, whereas the "in real life" image looks like it's at least six feet, based on it coming up over the top of the picture frame.
I ended up engaging a local woodworker to make me a piece like it instead. Obviously an order of magnitude difference in price, but it will actually be real solid walnut and finished to match my dining table.
gtowey•Jul 21, 2026
Even funnier is that first picture shows the bottom panel overlapping the baseboard in a way that is impossible in real life.
throwup238•Jul 21, 2026
The merchant is even called “VibesPlante”
jliptzin•Jul 21, 2026
If it doesn’t fit, then you must have gained weight while the item was in transit
agile-gift0262•Jul 21, 2026
You can fix it through our partner, Ozempic
jareklupinski•Jul 21, 2026
i'd like it to notice things i would miss, like "this is ring-spun shirt, so it will sit like this on your torso" or "these pleats will require you to iron them" etc
pwillia7•Jul 21, 2026
I bet if you had a way to easily train a LORA on you trying on various clothing types the models would do pretty well but I agree without that I can't think of a way to get it to work. Reminds me of the flattering mirrors scam https://finance.yahoo.com/news/company-swears-controversial-...
brookst•Jul 21, 2026
The more honest ones are “how it looks on you” and don’t promise fit that depends on so many measurements that aren’t even visible in pics.
Sure, it’s idealized, but some people benefit from seeing color / neckline / etc on themselves as a visual reference.
Me, I’m a text-learner so I don’t get it at all. But I know people who get value.
therealpygon•Jul 21, 2026
Good or bad, I’m pretty sure showing things in the best light has always been the point of “marketing”. There is a reason ads aren’t filled with ugly people with misshaped bodies, and it isn’t because the intention is to reflect reality, so what you allude to as problematic is what a lot of businesses call a feature.
sixothree•Jul 21, 2026
I'm pretty sure I saw somewhere people can sell re-sell clothes. They can include a picture of themselves wearing the item and it replaces literally everything but the clothing with someone/someplace "prettier".
In the end the one thing that is completely honest is the portion of the picture that is the item you are re-selling. But somehow to me the entire thing feels disingenuous.
therealpygon•Jul 21, 2026
Marketing a product is often disingenuous. You didn’t actually need that shirt, or that bag of chips. That’s the definition of successful marketing in their mind.
If marketing had to be honest, an entirely different set of products would likely be the most popular. I don’t think that’s a good thing by any means, just that the problem hasn’t really changed much beyond more focused targeting, which has always been a goal of marketing anyway.
It’s hard to convince someone who isn’t interested to buy something which is why tech ads sell on tech channels and car ads sell on car channels. Now you get to star in your very own clothes ads…for yourself…targeting you…
I suspect companies drool for that prospect because you are selling it to yourself with only some nice gentle nudging. That is the grail of marketing; that you don’t even realize you are in the midst of being marketed to.
LogicFailsMe•Jul 21, 2026
I use nanobanana 2 to test changes in paint and flooring/tiling to great success. And the real pro move is taking those images to a designer to tweak the remaining 20% or so.
sixothree•Jul 21, 2026
You just described how people are using AI in general. And I think that last step is actually a huge gap. People will brainstorm or improve their understanding of the task with AI and create something. Ideally they would hand it off to a professional for finalization like you did.
I wish a service existed where I could make something in AI (text, images, whatever), then pass that AI output to an actual human who would use it as a guide to produce an actual product.
paradox460•Jul 21, 2026
I used it recently to explore a few placements of a pool and glass building on my property. It's not authoritative, by any means, but it did answer some questions
And at this level it's barely any different than an architectural render, more for "how could this look" rather than "how will this look"
qwertox•Jul 21, 2026
> But these models will always make the clothes fit your body and show you in flattering light and so on
This is something that can be fixed over time. And if this forces clothing manufacturers to stick more to their advertised "specs" (width/length), then it's a win for us.
coldtea•Jul 21, 2026
Nobody that's pushing it has any incentive for fixing it - and they have all the opposite incentives.
victorbjorklund•Jul 21, 2026
Is it gonna be less distorted than just seeing the shirt on a model photographed by a professional in the perfect light?
collinmcnulty•Jul 21, 2026
Yes, because it is the actual dimensions of the real shirt.
smith7018•Jul 21, 2026
Many product/model shots actually use clips to make the clothing look like it's perfectly fitted to the body. [1] This image is actually over 15 years old at this point. I think there should be laws that prevent this because it veers into false advertising though others believe it's alright because you can theoretically tailor the clothes to fit like this.
Regardless, I'd rather see real clothing on a real person when it comes to my purchasing decisions. I buy a lot of vintage clothes online and I've noticed a dramatic uptick in AI images of models wearing the clothes. I've never once bought from those sellers because it feels disingenuous. Sometimes they have fake runways which is actual false advertising because it makes the item appear more expensive than it really is. I've also noticed that the AI models' body types are always thin even if the item is a L or XL. Needless to say, the AI isn't showing me what an XL looks like on a small model; it's showing what a small model would look like if the item fit perfectly.
That image isn't from a clothes catalog, they'd generally use the unaltered garment there.
phainopepla2•Jul 21, 2026
Nope. I used to assist a photographer who did shoots for catalogs, and clips were used quite a bit. The same model has to wear dozens of pieces of clothing throughout the shoot, and not everything is going to fit their body well, but the clients obviously want everything to look well fitted.
stavros•Jul 21, 2026
Ah, I stand corrected, thanks.
coldtea•Jul 21, 2026
Yes, because at least that won't be flattering you (not to mention it would be showing the actual garment).
spaceman_2020•Jul 21, 2026
A prompt went viral recently where people were sharing their pictures and asking chatgpt to visualise what their looksmatched partner would look like
The result was always someone extremely good looking
There’s going to be an entirely new class of mental disorders that will emerge from people being deluded by AI
zkmon•Jul 21, 2026
How does this fry pan look with a fish in it? Ask Mr Bean.
teraflop•Jul 21, 2026
The short-term goal of a tool like this is to sell products. The more ambitious long-term goal is to shift cultural norms, blurring the lines between advertising and reality until the question you're asking is no longer consciously asked. At least, not by average people, and not at the point of purchase.
I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.
dataflow•Jul 21, 2026
> I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy.
From what I hear (fact-check, etc. needed) that's already the case right now with much more consequential transactions, like renting real estate in NYC. Square footages that are blatant lies, etc.
verisimi•Jul 21, 2026
> I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy.
I find it easy now!
palmotea•Jul 21, 2026
> I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.
It's infuriating the amount of effort people will expend to claim that what was achieved in the past is literally impossible to do now. It's pervasive, especially from allegedly-smart people like software engineers.
dleeftink•Jul 21, 2026
Software has a dependency contract that is quite unique. We can't imagine the shelflife beyond the Bazaar or Cathedral. They seem ever lasting.
The myth of impossibility remains, because we can't imagine what it means to forever occupy those places, even when others rise past them.
ElProlactin•Jul 21, 2026
> I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy.
There hasn't been "truth in advertising" for many, many years. The only thing that has changed recently is that you don't even have to hide most of the lies.
supern0va•Jul 21, 2026
>I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.
Alternatively, perhaps we'll see models fine-tuned or steer-able towards accuracy that customers can themselves use to get a more honest view on what the product would look like in person.
The funny thing about these tools is that they can go either direction, but it sure seems like there's the potential for it empower individuals and shift the balance. Clothing sales shifting online has given sellers the advantage/ease to deceive without much customers can do other than hope the reviews aren't manipulated (they are) even before AI. Maybe this can turn things around as people start to shop with personal agents.
derefr•Jul 21, 2026
That's a fun thought. A gradual division between "institutional models" trained to have capabilities that solve for the needs of corporations, vs. "personal models" trained to have capabilities that solve for the needs of individuals.
I wonder what kind of capabilities those might be?
newswasboring•Jul 21, 2026
I've been struggling with this question myself. But isn't this a model training/use problem (a.k.a skill issue )? Isn't there a way to make these models be faithful to how people will actually look?
daedrdev•Jul 21, 2026
But they dont have to do either of these things. Thats just what people prompt
aljgz•Jul 21, 2026
One joy of online shopping, especially for people doing it in an impulsive and/or addictive way (it's not really rare) is the satisfaction they get from the imagination of having it.
An acquaintance of mine was buying many and not wearing most, as she did not attend that many social occasions. Still, she kept buying.
Eventually, she had to face the actual problem in her life that bothered her. She ended up dealing with it, terribly.
geniium•Jul 21, 2026
Yet it pushes the goal of the brand/shop forward: make the product more appealing and sell
sandworm101•Jul 21, 2026
User-targeted fashion advice is the polite goal. AI designed to generate images of people is actually racing to capture the entertainment markets. They want to be ready to replace models/actors in everything from fashion mags to porn studios. That is where the money is.
yorwba•Jul 21, 2026
For Alibaba, it's definitely all about shopping. That's where their money is. They haven't had much luck investing in entertainment so far.
jeromechoo•Jul 21, 2026
It’s clear to us here on HN but to the average person today the way these models actually work is beyond the realm of constraints and reasoning.
imhoguy•Jul 21, 2026
This will push us even more to go outside and visit actual shops. Becouse thanks to AI we will have more time? Will we?
viridir•Jul 21, 2026
This looks impressive!
pal9000i•Jul 21, 2026
How long until we get rid of the AI "plasticness" in portrait kind of generated images?
aitchnyu•Jul 21, 2026
There is an IRL phenomena, glass skin skincare and makeup, where the person's skin is evenly flat and toned and glossy. Did you think they are more plasticky than IRL models?
numpad0•Jul 21, 2026
That's supposed to make skin appear lively. Human skins in AI images tend to look clouded, opaque, and overall un-alive, so to speak.
jrs100000•Jul 21, 2026
The right models and LORAs can get rid of it right now. People apparently really like everyone to look like over exposed over filtered mannequins, so the big companies target that look.
cubefox•Jul 21, 2026
I think it's unintentional. People usually dislike any recognizable "AI look", but they do like other aspects which might have an unrealistic AI look as a side effect.
For example, Google's Imagen 3 usually looked a lot less fake than the newer Imagen 4, but the latter still scored higher on most benchmarks because it made fewer mistakes and had better prompt following capabilities.
A similar thing happened with Dalle-E 2 and Dell-E 3: The new model was better but also more fake looking.
lifeofpi331144•Jul 21, 2026
what are the right models and loras?
saltysalt•Jul 21, 2026
It will be interesting to compare this to Flux 2.
Oarch•Jul 21, 2026
I assume Van Gogh didn't paint enough hands to train from!
gpjanik•Jul 21, 2026
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
viraptor•Jul 21, 2026
This model isn't supposed to contain all the numerical data. It will give you a (usually) matching graph transformed from one you provide or from a table of information you provide. Or you can pipeline from an LLM doing research on that data first. But expecting an image gen model to get you GDP info has got to be one of the worst possible approaches.
> Especially text rendering
That's true though. I still got some completely fried letters in headings.
It included the table verbatim and even managed to hallucinate a reasonable heading for it, but then the graph doesn't even manage to align the data points with the time axis, leading to an unfortunate collision in the middle.
I guess you should use a traditional graphing library for your presentation slides for now.
spwa4•Jul 21, 2026
... but if you want an accurate graph, why not ask the LLM model to put the data points into a graphing library?
gpjanik•Jul 21, 2026
I am not expecting it, it's the Qwen team is claimingthey can do much harder tasks than this, like rendering a consistent page of a maths paper, or creating true to fact explainers.
They can't.
dsrtslnd23•Jul 21, 2026
Seems that this will not be open weights?
gchokov•Jul 21, 2026
It failed to create a simple overlay on a map - something ChatGPT had no issues with.
maxloh•Jul 21, 2026
I am curious whether the model requires a font to be installed. Does it also generate the glyphs for the text?
simonw•Jul 21, 2026
Yes, it generates the text without using a font. Same is true of other image models like ChatGPT Images and Gemini Nano Banana and Midjourney.
vunderba•Jul 21, 2026
It's a proprietary model as a service - it doesn't require anything outside of a browser. But even when they released open-weight (the original Qwen-Image) it's a diffusion model and can't use custom font files.
feverzsj•Jul 21, 2026
The "piss filter" is still everywhere.
dhbradshaw•Jul 21, 2026
The generated latex pdf!
hessammehr•Jul 21, 2026
Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
amrrs•Jul 21, 2026
that's a very interesting point to test new models. I speak an Indian language called "Tamil" and have always tested new models with Tamil but also have tried little bit of Arabic (quranic verses) with previous GPT image 2 and Nanobanana pro and they have nailed it. Don't know if it was because of extensive training data.
bejd•Jul 21, 2026
I wonder if they got permission to generate that (admittedly impressive) Berserk image.
arslan9063•Jul 21, 2026
THIS IS EXACTLY WHAT I WAS LOOKING FOR
ninjagoo•Jul 21, 2026
The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.
But: not open-source/open-weights, and no indication that weights/source will be released either.
simonw•Jul 21, 2026
> to precisely describe the full 3×3 grid takes a full 3.7k tokens
It's a shame they didn't share that prompt - it would make that demo more convincing.
tarcon•Jul 21, 2026
I am surprised by the rather bad output. It doesn't achieve qwen image 1 quality in composition or anatomical correctness. Tested on chat.qwen.ai
I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.
sajithdilshan•Jul 21, 2026
I truly wish these models were available when I was in University. As a visual learner it would have been much easier for me to understand certain topics with illustrative diagrams rather than reading a wall of text.
DaiPlusPlus•Jul 21, 2026
…but the illustrative diagrams are a simulacrum; if you ask Qwen, or any image-generator, for an “accurate” poster-design featuring a representation of a model of an atom and explaining its constituent parts I expect you’ll get an imitation-airbrush rendering of red, blue, and grey table-tennis balls orbiting in perfect circles; you might get an electron-shell diagram if you’re lucky. What you won’t get is anything remotely related to probability-clouds.
Edit: just to test myself I asked Nano Banana 2 to generate “an undergraduate infographic poster about how atoms work” - and the result was something right out of a middle-school science textbook and very Bohr…
sajithdilshan•Jul 21, 2026
This is the image I got for the same prompt: https://jumpshare.com/s/mqqBdl7U59FWPiEXjwoM. It's more like high-school level, but not bad. I can imagine a collage professor can improve the prompt the create a more accurate and detailed diagram
zarzavat•Jul 21, 2026
This seems like something that could be solved by asking an LLM to write the prompt for the image model. You can also feed in the output of an image model into an LLM and ask it to check it/make improvements.
R_D_Olivaw•Jul 21, 2026
This is the way.
There's definitely more dogfooding that needs to be done. And id argue that if your purpose is truly to learn or to teach, the process of describing that image will do wonders for retention.
DaiPlusPlus•Jul 21, 2026
> if your purpose is truly to learn or to teach, the process of describing that image will do wonders for retention.
If you're wanting to learn then you won't have the expertise to craft a prompt that has the correct details or to spot when the model makes a mistake in its output.
If you're in a teaching position then you won't have anything to learn.
verve_rat•Jul 21, 2026
Not necessarily. A reproduction task as a spaced repetition activity can help embed learning. If you are trying to write that prompt there is no reason you can't refer back to the text book to build the prompt and check the output.
wincy•Jul 21, 2026
My wife has taken all her recipes and fed them through ChatGPT image gen to make zine pages and they’re really cool! She’s building a recipe book for the kids so they’ll know all the recipes from their childhood.
lifthrasiir•Jul 21, 2026
> In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.
And yet the Korean text is not accurate... [1]
[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.
ndom91•Jul 21, 2026
Again not released on huggingface immediately?
timedude•Jul 21, 2026
Zooming in on mobile on that website causes a large white area to obstruct the page. Might wanna look into that.
As for the image model, wow...
Smaily•Jul 21, 2026
why I cant submitte new posts here ? My account since 2016
cubefox•Jul 21, 2026
Probably because you have too few upvoted comments. The rules likely got stricter when LLMs started to infiltrate Hacker News.
luciana1u•Jul 21, 2026
the natural endpoint of this technology is product photos that look better than the actual product, which is going to make unboxing videos the last remaining source of truth on the internet
jcattle•Jul 21, 2026
What I can not wrap my head around: How are these models trained?
What training mechanism or model architecture provides the glue to go from human text to images?
Don't you need to have millions of really descriptively labelled images?
There are ML models that do the reverse and output image to text, which assist quite a lot.
The better the text represents the unique thing in the photo, the better the model understands what that text means.
vonneumannstan•Jul 21, 2026
Short answer yes.
Slightly longer answer for older text to image models you teach them how to encode images and text into the same latent space. Then you simply do a conversion, take a text input, put it into latent space and then extract the image that latent space represents.
topheroo•Jul 21, 2026
I’d argue that talking about “authentic” AI-generated images is oxymoronic.
zzleeper•Jul 21, 2026
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
hawtads•Jul 21, 2026
Not sure about OCR specifically, but the newer (past quarter) vision language models all have a lot of post training on detecting garbled text specifically. You can feed some of the old stable diffusion outputs into a modern model and they can figure out pretty reliably if the text gen is mangled. I think the feature is probably used as part of the RL for the image gen to correct for bad text rendering.
geooff_•Jul 21, 2026
No pricing table? No benchmarks? This is just marketingslop.
Image token pricing has been fairly steady while text token prices fall, yet image model release discussion seems to be more focused on how beautiful the women the model generates are versus any sort of substantive discussion.
m3kw9•Jul 21, 2026
Impressive, but their woman face generation always use very similar, too perfect, same prettiness faces, it's very obvious.
dzonga•Jul 21, 2026
once a.i images took over - a lot of things that were based on images died.
going to a fancy restaurant coz of 'you can take good pics' - dead - a.i can recreate that cheaper.
which means for a certain demographic - dating apps are dead too - since those were largely based on swiping photos.
the premium of in-person / small intimate events has gone up. likewise meeting in person, or doing things with a person live.
this also means imperfection has gone up in value (imperfection is a human quality).
hankbond•Jul 21, 2026
the era of wabi-sabi is among us.
BoorishBears•Jul 21, 2026
"the premium of in person" is a string of words that only means something to people in a very specific circle where everyone is constantly trying to tastemake and write in lower case.
Ironically they have extremely limited influence on how the larger world moves: normal people won't meaningfully change their habits around eating good food or meeting up with each other just because AI can generate a picture of either.
jarjoura•Jul 21, 2026
I would attribute any in-person cultural shift to the generation of 20-somethings trapped online during their teenage years from COVID. They are finally old enough, and have enough power to run things.
Mr_Eri_Atlov•Jul 21, 2026
Image generation is the least interesting and most concerning aspect of AI.
I see lots of useful tools made related to translation and code generation, but image generation seems to be primarily used to deceive, harass, or embelish to the point of questioning what the point of a photo is anymore anyway?
kthartic•Jul 21, 2026
I can give one innocent example - I made an app that summarises D&D sessions, including images to depict some of the scenes.
But I understand the jaded perspective. I guess the idea of ‘deepfakes’ didn’t give AI image gen a great reputational start.
lalith_c•Jul 21, 2026
didn’t anthropic accuse them of stealing fable?
flakiness•Jul 21, 2026
It's a bit shocking to see them showing off blatantly disinformation-al examples. See the Japanese manga one. It says "原作 監修 三浦建太郎" which says its original is written by, and it is supervised the by, the famous author (who died a few years ago so who can he possibly supervise?) and "共同制作 白泉社" saying: collaborated by a (famous Japanese) publisher.
I hope these stakeholders have good enough layers.
noodlescb•Jul 21, 2026
It's hilarious how bad image generation is. I just tried this out and I gave it two actual logos, a full design language spec, and three screenshots of the actual application and it spat out three absolutely awful images, none of which had the same logo as the one I sent.
43 Comments
Impressive.
Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?
Any suggestions for the best open, non-opinionated model?
https://pastes.io/uenL6X9K
It also seems to have an obsession with this celebrity, based on how many times ctrl-f for "stefani" turns up a result.
https://en.wikipedia.org/wiki/Gwen_Stefani
Do you want to know more?
[ ] Yes [x] no
erm... what?
[whynotboth.gif]
Am i gregnant?
https://youtu.be/EShUeudtaFg
https://knowyourmeme.com/memes/bobs-and-vegana
> Please don't post comments saying that HN is turning into Reddit. It's a semi-noob illusion, as old as the hills.
This is pathologically outside scope. I don't think I have ever seen somebody threaten somebody else on hn before. The 'don't make me do it to you' abuser trope is next level.
(See kids, it is possible to fight memes/racism with memes! And well... yeah, this really is racist.)
That said it's possible search engines in China or other countries might use it, but it's very easy to game so it doesn't really make sense
https://www.nytimes.com/2019/06/07/us/hate-groups-porn-consp...
porn is as American as apple pie.
Apart from that, I made no judgement about porn in general, just about porn tags in Chinese backed AI websites.
So not at all then as Apple Pie is a traditional English desert. :P
https://en.wikipedia.org/wiki/Apple_pie
Run this in console to see all the tags:
(i.e. the porn references)
If you look at past snapshots at archive.org, you notice that the meta keywords are growing like an append-only list, which means it's probably part of some messed up seo pipeline. The other clue is that it's stuffing the meta keywords which apparently only Yandex uses as a search signal[0].
The list has 3882 entries. A lot of them are clustered and look like auto-complete results. But a bunch of them look like hyper-specific, misspelled search results (e.g. "145 gwen rd cheshire ct"). Google's webmaster tools doesn't provide distinct queries like that, but yandex's does[1].
My best guess of what's happening is that Qwen is monitoring it's search queries in Yandex, dumping that list into a serp service that scrapes yandex's autocomplete suggestions, and then taking that list and dumping it into their meta keywords.
It explains the urls in the list (ppls using search engines like address bars), the seeming fixation around certain topics (which usually starts with a misspelling), and the random one-off queries.
0: https://yandex.com/support/webmaster/en/controlling-robot/me...
1: https://yandex.com/support/webmaster/en/service/popular-quer...
It apparently adds common search terms that contain words like "qwen". This evidently includes possibly mistyped searches for "gwen" or "ben" in a NSFW context.
Maybe someone knows more about how such SEO tools work, and where they pull the data from.
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle
When I want to emphasize something, I tend to repeat it
Results are mixed, expensive, but it really feels you're few months off the next improvement to really nail it. It's already good enough.
Wonder what Qwen image will provide over nano banana.
Here's a recent example: https://www.etsy.com/ca/listing/4509158065/corner-wall-shelf...
Ironically, ChatGPT is decently good at ferreting these out. Like I sent it a screenshot of that listing and it not only helped me find where the original item was for sale, but also pointed out how the dimensioned diagram shows it as being just 49" tall, whereas the "in real life" image looks like it's at least six feet, based on it coming up over the top of the picture frame.
I ended up engaging a local woodworker to make me a piece like it instead. Obviously an order of magnitude difference in price, but it will actually be real solid walnut and finished to match my dining table.
Sure, it’s idealized, but some people benefit from seeing color / neckline / etc on themselves as a visual reference.
Me, I’m a text-learner so I don’t get it at all. But I know people who get value.
In the end the one thing that is completely honest is the portion of the picture that is the item you are re-selling. But somehow to me the entire thing feels disingenuous.
If marketing had to be honest, an entirely different set of products would likely be the most popular. I don’t think that’s a good thing by any means, just that the problem hasn’t really changed much beyond more focused targeting, which has always been a goal of marketing anyway.
It’s hard to convince someone who isn’t interested to buy something which is why tech ads sell on tech channels and car ads sell on car channels. Now you get to star in your very own clothes ads…for yourself…targeting you…
I suspect companies drool for that prospect because you are selling it to yourself with only some nice gentle nudging. That is the grail of marketing; that you don’t even realize you are in the midst of being marketed to.
I wish a service existed where I could make something in AI (text, images, whatever), then pass that AI output to an actual human who would use it as a guide to produce an actual product.
And at this level it's barely any different than an architectural render, more for "how could this look" rather than "how will this look"
This is something that can be fixed over time. And if this forces clothing manufacturers to stick more to their advertised "specs" (width/length), then it's a win for us.
Regardless, I'd rather see real clothing on a real person when it comes to my purchasing decisions. I buy a lot of vintage clothes online and I've noticed a dramatic uptick in AI images of models wearing the clothes. I've never once bought from those sellers because it feels disingenuous. Sometimes they have fake runways which is actual false advertising because it makes the item appear more expensive than it really is. I've also noticed that the AI models' body types are always thin even if the item is a L or XL. Needless to say, the AI isn't showing me what an XL looks like on a small model; it's showing what a small model would look like if the item fit perfectly.
[1] https://www.primermagazine.com/wp-content/uploads/2011/02/St...
The result was always someone extremely good looking
There’s going to be an entirely new class of mental disorders that will emerge from people being deluded by AI
I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.
From what I hear (fact-check, etc. needed) that's already the case right now with much more consequential transactions, like renting real estate in NYC. Square footages that are blatant lies, etc.
I find it easy now!
It's infuriating the amount of effort people will expend to claim that what was achieved in the past is literally impossible to do now. It's pervasive, especially from allegedly-smart people like software engineers.
The myth of impossibility remains, because we can't imagine what it means to forever occupy those places, even when others rise past them.
There hasn't been "truth in advertising" for many, many years. The only thing that has changed recently is that you don't even have to hide most of the lies.
Alternatively, perhaps we'll see models fine-tuned or steer-able towards accuracy that customers can themselves use to get a more honest view on what the product would look like in person.
The funny thing about these tools is that they can go either direction, but it sure seems like there's the potential for it empower individuals and shift the balance. Clothing sales shifting online has given sellers the advantage/ease to deceive without much customers can do other than hope the reviews aren't manipulated (they are) even before AI. Maybe this can turn things around as people start to shop with personal agents.
I wonder what kind of capabilities those might be?
An acquaintance of mine was buying many and not wearing most, as she did not attend that many social occasions. Still, she kept buying.
Eventually, she had to face the actual problem in her life that bothered her. She ended up dealing with it, terribly.
For example, Google's Imagen 3 usually looked a lot less fake than the newer Imagen 4, but the latter still scored higher on most benchmarks because it made fewer mistakes and had better prompt following capabilities.
A similar thing happened with Dalle-E 2 and Dell-E 3: The new model was better but also more fake looking.
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
> Especially text rendering
That's true though. I still got some completely fried letters in headings.
It included the table verbatim and even managed to hallucinate a reasonable heading for it, but then the graph doesn't even manage to align the data points with the time axis, leading to an unfortunate collision in the middle.
I guess you should use a traditional graphing library for your presentation slides for now.
They can't.
But: not open-source/open-weights, and no indication that weights/source will be released either.
It's a shame they didn't share that prompt - it would make that demo more convincing.
I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.
Edit: just to test myself I asked Nano Banana 2 to generate “an undergraduate infographic poster about how atoms work” - and the result was something right out of a middle-school science textbook and very Bohr…
There's definitely more dogfooding that needs to be done. And id argue that if your purpose is truly to learn or to teach, the process of describing that image will do wonders for retention.
If you're wanting to learn then you won't have the expertise to craft a prompt that has the correct details or to spot when the model makes a mistake in its output.
If you're in a teaching position then you won't have anything to learn.
And yet the Korean text is not accurate... [1]
[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.
As for the image model, wow...
What training mechanism or model architecture provides the glue to go from human text to images?
Don't you need to have millions of really descriptively labelled images?
There are ML models that do the reverse and output image to text, which assist quite a lot.
The better the text represents the unique thing in the photo, the better the model understands what that text means.
Slightly longer answer for older text to image models you teach them how to encode images and text into the same latent space. Then you simply do a conversion, take a text input, put it into latent space and then extract the image that latent space represents.
Image token pricing has been fairly steady while text token prices fall, yet image model release discussion seems to be more focused on how beautiful the women the model generates are versus any sort of substantive discussion.
going to a fancy restaurant coz of 'you can take good pics' - dead - a.i can recreate that cheaper.
which means for a certain demographic - dating apps are dead too - since those were largely based on swiping photos.
the premium of in-person / small intimate events has gone up. likewise meeting in person, or doing things with a person live.
this also means imperfection has gone up in value (imperfection is a human quality).
Ironically they have extremely limited influence on how the larger world moves: normal people won't meaningfully change their habits around eating good food or meeting up with each other just because AI can generate a picture of either.
I see lots of useful tools made related to translation and code generation, but image generation seems to be primarily used to deceive, harass, or embelish to the point of questioning what the point of a photo is anymore anyway?
But I understand the jaded perspective. I guess the idea of ‘deepfakes’ didn’t give AI image gen a great reputational start.
I hope these stakeholders have good enough layers.