Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge
371 points
6 hours ago
| 36 comments
| qwen.ai
| HN
mynti
5 hours ago
[-]
To me it feels so weird that people are trying to push these model for online shopping like "here is how this dress/shirt/pants would look on you". But these models will always make the clothes fit your body and show you in flattering light and so on. How the actual garment fits is still as elusive as before these tools
reply
teraflop
34 minutes ago
[-]
The short-term goal of a tool like this is to sell products. The more ambitious long-term goal is to shift cultural norms, blurring the lines between advertising and reality until the question you're asking is no longer consciously asked. At least, not by average people, and not at the point of purchase.

I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.

reply
dataflow
10 minutes ago
[-]
> I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy.

From what I hear that's already the case right now with much more consequential transactions, like renting real estate in NYC. Square footages that are blatant lies, etc.

reply
verisimi
5 minutes ago
[-]
> I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy.

I find it easy now!

reply
khurs
4 hours ago
[-]
As long as the Product Manger gets a bonus/promotion then it's a success.
reply
victorbjorklund
1 hour ago
[-]
Is it gonna be less distorted than just seeing the shirt on a model photographed by a professional in the perfect light?
reply
collinmcnulty
1 hour ago
[-]
Yes, because it is the actual dimensions of the real shirt.
reply
smith7018
1 hour ago
[-]
Many product/model shots actually use clips to make the clothing look like it's perfectly fitted to the body. [1] This image is actually over 15 years old at this point. I think there should be laws that prevent this because it veers into false advertising though others believe it's alright because you can theoretically tailor the clothes to fit like this.

Regardless, I'd rather see real clothing on a real person when it comes to my purchasing decisions. I buy a lot of vintage clothes online and I've noticed a dramatic uptick in AI images of models wearing the clothes. I've never once bought from those sellers because it feels disingenuous. Sometimes they have fake runways which is actual false advertising because it makes the item appear more expensive than it really is. I've also noticed that the AI models' body types are always thin even if the item is a L or XL. Needless to say, the AI isn't showing me what an XL looks like on a small model; it's showing what a small model would look like if the item fit perfectly.

[1] https://www.primermagazine.com/wp-content/uploads/2011/02/St...

reply
coldtea
1 hour ago
[-]
Yes, because at least that won't be flattering you (not to mention it would be showing the actual garment).
reply
k2enemy
3 hours ago
[-]
I've noticed something similar in Facebook marketplace ads for used furniture. Most of the images are AI generated to look like a Pottery Barn catalog, then the last image will be the actual item, full of scratches and other damage, sitting in a messy garage.
reply
mikepurvis
59 minutes ago
[-]
I've started to see this on Etsy and Wayfair too, where there will be a listing that is clearly just MDF flatpack being resold from China, but the AI-generated images wildly exaggerate the proportions of it.

Here's a recent example: https://www.etsy.com/ca/listing/4509158065/corner-wall-shelf...

Ironically, ChatGPT is decently good at ferreting these out. Like I sent it a screenshot of that listing and it not only helped me find where the original item was for sale, but also pointed out how the dimensioned diagram shows it as being just 49" tall, whereas the "in real life" image looks like it's at least six feet, based on it coming up over the top of the picture frame.

I ended up engaging a local woodworker to make me a piece like it instead.

reply
aembleton
3 hours ago
[-]
Estate agents are also doing it; making interiors of houses look very different to how they really are.
reply
spaceman_2020
1 hour ago
[-]
This should really be illegal
reply
jeremyjh
2 hours ago
[-]
Yes and even if all they do is ask it to stage an empty room picture with furniture and decorations it will make improvements like adding a doorway to a room that doesn’t exist.
reply
agile-gift0262
2 hours ago
[-]
The worst I've seen changed the view out a window from an alley to a private garden. Most also add natural light that isn't possible, and in some cases I'm convinced the generate image is depicting the space as larger than it actually is, with furniture and spaces that wouldn't fit in reality
reply
hiccuphippo
1 hour ago
[-]
I once heard about a company that specialized in making furniture about 10% smaller than normal for use in show rooms to make the spaces look larger. Not surprised they do it with AI if they can.
reply
masfuerte
8 minutes ago
[-]
This is a thing in show homes on new-build estates in the UK. British houses are often tiny.
reply
mikepurvis
57 minutes ago
[-]
Wide angle lenses have been a thing in real estate photography since forever, but you would always have the reference of the furniture to ground your perception. Having furniture be slightly downsized is diabolical.
reply
skinfaxi
2 hours ago
[-]
Seems like false advertising.
reply
QuantumGood
3 hours ago
[-]
And knocking out the messy garage background to replace it with a showroom is the easiest prompt.
reply
number6
4 hours ago
[-]
Sounds exactly like something the marketing department would use to up the convertion rate
reply
brookst
2 hours ago
[-]
The more honest ones are “how it looks on you” and don’t promise fit that depends on so many measurements that aren’t even visible in pics.

Sure, it’s idealized, but some people benefit from seeing color / neckline / etc on themselves as a visual reference.

Me, I’m a text-learner so I don’t get it at all. But I know people who get value.

reply
viraptor
4 hours ago
[-]
There's definitely a predatory "everything will look good", but you can also leverage exactly that part for your own purposes - I've done pre-shopping a couple of times by asking for something like "a grid of 9 versions of my photo, wearing different types of X that look good". Definitely helped with choosing a good style.
reply
pwillia7
2 hours ago
[-]
I bet if you had a way to easily train a LORA on you trying on various clothing types the models would do pretty well but I agree without that I can't think of a way to get it to work. Reminds me of the flattering mirrors scam https://finance.yahoo.com/news/company-swears-controversial-...
reply
therealpygon
2 hours ago
[-]
Good or bad, I’m pretty sure showing things in the best light has always been the point of “marketing”. There is a reason ads aren’t filled with ugly people with misshaped bodies, and it isn’t because the intention is to reflect reality, so what you allude to as problematic is what a lot of businesses call a feature.
reply
sixothree
11 minutes ago
[-]
I'm pretty sure I saw somewhere people can sell re-sell clothes. They can include a picture of themselves wearing the item and it replaces literally everything but the clothing with someone/someplace "prettier".

In the end the one thing that is completely honest is the portion of the picture that is the item you are re-selling. But somehow to me the entire thing feels disingenuous.

reply
jliptzin
3 hours ago
[-]
If it doesn’t fit, then you must have gained weight while the item was in transit
reply
agile-gift0262
2 hours ago
[-]
You can fix it through our partner, Ozempic
reply
zkmon
53 minutes ago
[-]
How does this fry pan look with a fish in it? Ask Mr Bean.
reply
LogicFailsMe
2 hours ago
[-]
I use nanobanana 2 to test changes in paint and flooring/tiling to great success. And the real pro move is taking those images to a designer to tweak the remaining 20% or so.
reply
walrus01
5 hours ago
[-]
The cynic in me says this is just an advancement and logical continuation from the known problem of sycophantic behavior in text-to-text chatbot format LLM, to image generation models.
reply
inerte
2 hours ago
[-]
To me the tools are newer, but fit people in a controlled context to sell clothing is as old as photography itself.
reply
jareklupinski
3 hours ago
[-]
i'd like it to notice things i would miss, like "this is ring-spun shirt, so it will sit like this on your torso" or "these pleats will require you to iron them" etc
reply
spaceman_2020
1 hour ago
[-]
A prompt went viral recently where people were sharing their pictures and asking chatgpt to visualise what their looksmatched partner would look like

The result was always someone extremely good looking

There’s going to be an entirely new class of mental disorders that will emerge from people being deluded by AI

reply
qwertox
2 hours ago
[-]
> But these models will always make the clothes fit your body and show you in flattering light and so on

This is something that can be fixed over time. And if this forces clothing manufacturers to stick more to their advertised "specs" (width/length), then it's a win for us.

reply
coldtea
1 hour ago
[-]
Nobody that's pushing it has any incentive for fixing it - and they have all the opposite incentives.
reply
newswasboring
29 minutes ago
[-]
I've been struggling with this question myself. But isn't this a model training/use problem (a.k.a skill issue )? Isn't there a way to make these models be faithful to how people will actually look?
reply
morningsam
3 hours ago
[-]
Maybe this will lead to more business for tailors doing alterations, assuming the clothes people end up buying are expensive enough to justify it.
reply
jimnotgym
3 hours ago
[-]
I understand they are already doing decent business since the advent of the magic weight loss pen
reply
TylerE
4 hours ago
[-]
To be fair in many cases the actual clothes aren't much better. I've had two pair of the same pants, same brand, same size fit noticeably differently.
reply
epolanski
3 hours ago
[-]
I have a similar use case at work for previewing construction material and such in our catalogue applied to user uploaded images.

Results are mixed, expensive, but it really feels you're few months off the next improvement to really nail it. It's already good enough.

Wonder what Qwen image will provide over nano banana.

reply
weird-eye-issue
5 hours ago
[-]
The meta keywords in the HTML is very interesting. 100+ references to NSFW topics such as hentai, nudes, etc.
reply
01284a7e
7 minutes ago
[-]
What value does the 'keywords' meta tag even have these days? 77+ KB of crap added to the page weight. Web development is full of idiots.
reply
postalcoder
5 hours ago
[-]
That's hilarious. These keywords are applied globally, even on pages like https://qwen.ai/usagepolicy, where they very kindly ask you not to use their products for sexual content.

Run this in console to see all the tags:

  document.querySelector('meta[name="keywords"]').content
reply
khurs
4 hours ago
[-]
WTF lol

    who made qwen stefani's dress

    spiderman into the spider verse did qwen meet Peter

    qwen stacey porn

    do blake shelton and qwen stafani have children together
reply
OkWing99
3 hours ago
[-]
Some of it is a reference to https://en.wikipedia.org/wiki/Spider-Woman_(Gwen_Stacy)

(i.e. the porn references)

reply
drcongo
1 hour ago
[-]

    ben 10 four arms porn qwen
These keywords are kinda wild.
reply
Schlagbohrer
17 minutes ago
[-]
Which LLM do you think they used to generate the keywords
reply
LoveMortuus
3 hours ago
[-]
Why would they added URLs for porn websites to their keywords? Like literally the names of websites with ".com" and such at the end.
reply
user43928
4 hours ago
[-]
I assume they use a SEO tool that automatically adds these meta keywords to optimize for some search engines.

It apparently adds common search terms that contain words like "qwen". This evidently includes possibly mistyped searches for "gwen" or "ben" in a NSFW context.

Maybe someone knows more about how such SEO tools work, and where they pull the data from.

reply
spiderfarmer
3 hours ago
[-]
Sounds like “keyword shitter” (genuine tool)
reply
chmod775
5 hours ago
[-]
Bonus points for also covering misspellings like "pregbant".
reply
QuantumNomad_
5 hours ago
[-]
how is prangent formed

Am i gregnant?

https://youtu.be/EShUeudtaFg

reply
Geee
2 hours ago
[-]
how is babby formed? how girl get pragnent?
reply
user_7832
1 hour ago
[-]
If a women has starch masks on her body does that mean she has been pargnet before.?
reply
jmuguy
2 hours ago
[-]
mothers who kill thier babby. becuse these babbys cant frigth back?
reply
yewenjie
2 hours ago
[-]
how do i stop my son from looking at pikachu porn
reply
timedude
2 hours ago
[-]
Stork bring baby
reply
Mtinie
3 hours ago
[-]
qregnant
reply
gchamonlive
3 hours ago
[-]
Can you even get pregnato
reply
isoprophlex
5 hours ago
[-]
show bobs and vegana!
reply
user_7832
1 hour ago
[-]
That's racist!

(See kids, it is possible to fight memes/racism with memes! And well... yeah, this really is racist.)

reply
walrus01
5 hours ago
[-]
reply
marginalia_nu
2 hours ago
[-]
This reddit slop is unbecoming of HN.
reply
gchamonlive
2 hours ago
[-]
Is there such a thing as Reddit "unslop"? Thought it went without saying.
reply
hedora
55 minutes ago
[-]
Please don’t violate community guidelines like this. All HN threads look like this. Don’t make me report you to a submod.

> Please don't post comments saying that HN is turning into Reddit. It's a semi-noob illusion, as old as the hills.

reply
sgc
25 minutes ago
[-]
> Don’t make me report you to a submod.

This is pathologically outside scope. I don't think I have ever seen somebody threaten somebody else on hn before. The 'don't make me do it to you' abuser trope is next level.

reply
walrus01
5 hours ago
[-]
Wow you really weren't kidding. Fully automated AI slop tentacles for everyone!

https://pastes.io/uenL6X9K

It also seems to have an obsession with this celebrity, based on how many times ctrl-f for "stefani" turns up a result.

https://en.wikipedia.org/wiki/Gwen_Stefani

reply
goobatrooba
3 hours ago
[-]
Wow, crazy things - "shemale porn", "spider fucks venom", and various porn URLs like tushy.com. What the hell did they do to their AI? Is that the training corpus or intended use case?
reply
walrus01
3 hours ago
[-]
> Is that the training corpus or intended use case?

[whynotboth.gif]

reply
weird-eye-issue
5 hours ago
[-]
> spider qwen thicc, spider qwen tied up, spider qwen toddler costume, spider qwen trans, spider qwen wallpaper, spider qwen xxx hentai
reply
walrus01
5 hours ago
[-]
spiderman pointing at spiderman meme trans wallpaper xxx thicc hentai
reply
Mashimo
5 hours ago
[-]
> total drama qwen hentai

Do you want to know more?

[ ] Yes [x] no

reply
jpfromlondon
5 hours ago
[-]
Qwentai
reply
WithinReason
5 hours ago
[-]
probably due to misspellings as qwen stefani
reply
pdpi
4 hours ago
[-]
Keyword typosquatting I guess?
reply
alex_suzuki
5 hours ago
[-]
> ben 10 hentai game where qwen gets drunk

erm... what?

reply
archon1410
4 hours ago
[-]
Surely it was meant to be "Gwen" (the protagonist's cousin and the other main character in the show). I guess it was somehow (incorrectly) assumed to be a misspelling of Qwen and included in the tags. Or perhaps many people were misspelling Gwen as "Qwen" and it was all hoovered up.
reply
walrus01
5 hours ago
[-]
> qwen ten hentai, qwen ten xxx, qwen tennason footjob, qwen tennison hentai, qwen tennyson, qwen tennyson ass expansion, qwen tennyson giantess pussy, qwen tennyson hentai, qwen tennyson porn, qwen tft, qwen the milk maid, qwen the milkmaid, qwen the parts over, qwen the wolf witcher, qwen tire, qwen tneyson porn, qwen to go feelies youtube, qwen tokenizer java,
reply
bbor
5 hours ago
[-]
Whelp today’s a unique day on hacker news, wow! Didn’t expect to read that at my desk lol
reply
lukan
5 hours ago
[-]
That is hilarious as porn is illegal in China. But I guess pron SEO is allright, if it is against the west.
reply
weird-eye-issue
5 hours ago
[-]
Western search engines simply ignore the meta keywords tag for over a decade now
reply
lukan
5 hours ago
[-]
So what is the point then?
reply
weird-eye-issue
4 hours ago
[-]
There is no point. It's from people that have no clue which is most people that do search engine optimization

That said it's possible search engines in China or other countries might use it, but it's very easy to game so it doesn't really make sense

reply
vitalyan8184
4 hours ago
[-]
uh, are you implying that porn is bad somehow? I can show you a hundred articles that say only nazi incel chuds believe so.

https://www.nytimes.com/2019/06/07/us/hate-groups-porn-consp...

porn is as American as apple pie.

reply
Lio
1 hour ago
[-]
> porn is as American as apple pie

So not at all then as Apple Pie is a traditional English desert. :P

https://en.wikipedia.org/wiki/Apple_pie

reply
lukan
3 hours ago
[-]
Incels are probably the biggest porn consumers.

Apart from that, I made no judgement about porn in general, just about porn tags in Chinese backed AI websites.

reply
WhereIsTheTruth
4 hours ago
[-]
Brainrot is the new opium war
reply
cyanydeez
4 hours ago
[-]
how the turn tables
reply
postalcoder
5 hours ago
[-]
reply
yorwba
5 hours ago
[-]
GPT Image 1 ended up with a yellow tint without training on another image-generation model's output. It's just that humans like pictures with a soft sunset glow, and this is a very easy global signal for a preference model to pick up on, and for a image-generation model to imitate. So optimizing for aesthetic appeal makes everything slightly tinted by default, unless you make sure to countersteer.
reply
postalcoder
5 hours ago
[-]
Fair point!
reply
Grimblewald
5 hours ago
[-]
Didnt the piss tint come after their big studio ghibli heist?
reply
hugmynutus
1 hour ago
[-]
yellow/red tint is an extremely common problem not matter the photograph source you train on

source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle

reply
dannyw
4 hours ago
[-]
AI-generated images are part of the web now, if you're doing ordinary web scraping, you can't avoid training on generated images.
reply
greyb
26 minutes ago
[-]
You clearly haven't met a Chinese RedNote user.
reply
zzleeper
15 minutes ago
[-]
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
reply
jcattle
35 minutes ago
[-]
What I can not wrap my head around: How are these models trained?

What training mechanism or model architecture provides the glue to go from human text to images?

Don't you need to have millions of really descriptively labelled images?

reply
nucleative
29 minutes ago
[-]
That's exactly how they do it.

There are ML models that do the reverse and output image to text, which assist quite a lot.

The better the text represents the unique thing in the photo, the better the model understands what that text means.

reply
vonneumannstan
11 minutes ago
[-]
Short answer yes.

Slightly longer answer for older text to image models you teach them how to encode images and text into the same latent space. Then you simply do a conversion, take a text input, put it into latent space and then extract the image that latent space represents.

reply
mistercheph
32 minutes ago
[-]
I found this 3b1b guest video on diffusion helpful: https://www.youtube.com/watch?v=iv-5mZ_9CPY
reply
hessammehr
3 hours ago
[-]
Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
reply
amrrs
3 hours ago
[-]
that's a very interesting point to test new models. I speak an Indian language called "Tamil" and have always tested new models with Tamil but also have tried little bit of Arabic (quranic verses) with previous GPT image 2 and Nanobanana pro and they have nailed it. Don't know if it was because of extensive training data.
reply
topheroo
25 minutes ago
[-]
I’d argue that talking about “authentic” AI-generated images is oxymoronic.
reply
simonw
3 hours ago
[-]
> to precisely describe the full 3×3 grid takes a full 3.7k tokens

It's a shame they didn't share that prompt - it would make that demo more convincing.

reply
tarcon
3 hours ago
[-]
I am surprised by the rather bad output. It doesn't achieve qwen image 1 quality in composition or anatomical correctness. Tested on chat.qwen.ai

I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.

reply
sajithdilshan
3 hours ago
[-]
I truly wish these models were available when I was in University. As a visual learner it would have been much easier for me to understand certain topics with illustrative diagrams rather than reading a wall of text.
reply
DaiPlusPlus
2 hours ago
[-]
…but the illustrative diagrams are a simulacrum; if you ask Qwen, or any image-generator, for an “accurate” poster-design featuring a representation of a model of an atom and explaining its constituent parts I expect you’ll get an imitation-airbrush rendering of red, blue, and grey table-tennis balls orbiting in perfect circles; you might get an electron-shell diagram if you’re lucky. What you won’t get is anything remotely related to probability-clouds.

Edit: just to test myself I asked Nano Banana 2 to generate “an undergraduate infographic poster about how atoms work” - and the result was something right out of a middle-school science textbook and very Bohr…

reply
sajithdilshan
2 hours ago
[-]
This is the image I got for the same prompt: https://jumpshare.com/s/mqqBdl7U59FWPiEXjwoM. It's more like high-school level, but not bad. I can imagine a collage professor can improve the prompt the create a more accurate and detailed diagram
reply
zarzavat
2 hours ago
[-]
This seems like something that could be solved by asking an LLM to write the prompt for the image model. You can also feed in the output of an image model into an LLM and ask it to check it/make improvements.
reply
wincy
35 minutes ago
[-]
My wife has taken all her recipes and fed them through ChatGPT image gen to make zine pages and they’re really cool! She’s building a recipe book for the kids so they’ll know all the recipes from their childhood.
reply
embedding-shape
6 hours ago
[-]
Not a single word about when/if they'll actually release the weights for this, or am I missing it somewhere?
reply
user43928
5 hours ago
[-]
Considering Qwen-Image-2.0 weights have not been released either, it unfortunately looks unlikely.
reply
vachina
4 hours ago
[-]
Why release it just so Cursor/Azure/Amazon can profit off of it? Unless OpenAI actually opens it up fat chance.
reply
feverzsj
3 hours ago
[-]
The "piss filter" is still everywhere.
reply
gchokov
4 hours ago
[-]
It failed to create a simple overlay on a map - something ChatGPT had no issues with.
reply
Mashimo
5 hours ago
[-]
> Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.

Impressive.

Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?

reply
woadwarrior01
5 hours ago
[-]
Krea-2-Turbo. I've even got it working locally on my M5 iPad Pro.
reply
Izmaki
3 hours ago
[-]
> We implemented safety measures across the full model development lifecycle.

Any suggestions for the best open, non-opinionated model?

reply
woadwarrior01
2 hours ago
[-]
It is a reasonably non-opinionated model. My usual test is to ask these models to generate comic book and cartoon characters that hosted image generation models refuse to generate. I think that text is just CYA legalese.
reply
bitexploder
2 hours ago
[-]
You can also ablate it.
reply
Mashimo
5 hours ago
[-]
thanks mate. Sadly not supported yet by Invoke, but I will take a look.
reply
ninjagoo
3 hours ago
[-]
The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.

But: not open-source/open-weights, and no indication that weights/source will be released either.

reply
timedude
2 hours ago
[-]
Zooming in on mobile on that website causes a large white area to obstruct the page. Might wanna look into that.

As for the image model, wow...

reply
pal9000i
5 hours ago
[-]
How long until we get rid of the AI "plasticness" in portrait kind of generated images?
reply
jrs100000
4 hours ago
[-]
The right models and LORAs can get rid of it right now. People apparently really like everyone to look like over exposed over filtered mannequins, so the big companies target that look.
reply
aitchnyu
4 hours ago
[-]
There is an IRL phenomena, glass skin skincare and makeup, where the person's skin is evenly flat and toned and glossy. Did you think they are more plasticky than IRL models?
reply
numpad0
3 hours ago
[-]
That's supposed to make skin appear lively. Human skins in AI images tend to look clouded, opaque, and overall un-alive, so to speak.
reply
dsrtslnd23
4 hours ago
[-]
Seems that this will not be open weights?
reply
Oarch
5 hours ago
[-]
I assume Van Gogh didn't paint enough hands to train from!
reply
lifthrasiir
2 hours ago
[-]
> In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.

And yet the Korean text is not accurate... [1]

[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.

reply
luciana1u
1 hour ago
[-]
the natural endpoint of this technology is product photos that look better than the actual product, which is going to make unboxing videos the last remaining source of truth on the internet
reply
bejd
3 hours ago
[-]
I wonder if they got permission to generate that (admittedly impressive) Berserk image.
reply
ndom91
2 hours ago
[-]
Again not released on huggingface immediately?
reply
maxloh
4 hours ago
[-]
I am curious whether the model requires a font to be installed. Does it also generate the glyphs for the text?
reply
simonw
3 hours ago
[-]
Yes, it generates the text without using a font. Same is true of other image models like ChatGPT Images and Gemini Nano Banana and Midjourney.
reply
dhbradshaw
3 hours ago
[-]
The generated latex pdf!
reply
spwa4
5 hours ago
[-]
Appears to be closed-weights entirely. No word at all on any weights release.
reply
saltysalt
5 hours ago
[-]
It will be interesting to compare this to Flux 2.
reply
treetalker
5 hours ago
[-]
The red-dress woman's vestigial pinkie toes …
reply
xiaoyu2006
6 hours ago
[-]
The blog write-up style is so casual haha.
reply
jdw64
5 hours ago
[-]
Wow, it displays Korean properly without breaking. But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
reply
sheept
5 hours ago
[-]
I find it mildly interesting how your comment repeats itself but with different phrasing
reply
jdw64
4 hours ago
[-]
It's because of the structure of Korean. I think in Korean first, so I end up translating it directly.

When I want to emphasize something, I tend to repeat it

reply
rvz
5 hours ago
[-]
Midjourney already knew that image generation was going to zero. Again yet another reason why the model was never a moat in the first place.
reply
mahimai
5 hours ago
[-]
interesting
reply
viridir
5 hours ago
[-]
This looks impressive!
reply
arslan9063
3 hours ago
[-]
THIS IS EXACTLY WHAT I WAS LOOKING FOR
reply
gpjanik
4 hours ago
[-]
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.

Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.

reply
viraptor
4 hours ago
[-]
This model isn't supposed to contain all the numerical data. It will give you a (usually) matching graph transformed from one you provide or from a table of information you provide. Or you can pipeline from an LLM doing research on that data first. But expecting an image gen model to get you GDP info has got to be one of the worst possible approaches.

> Especially text rendering

That's true though. I still got some completely fried letters in headings.

reply
gpjanik
1 hour ago
[-]
I am not expecting it, it's the Qwen team is claimingthey can do much harder tasks than this, like rendering a consistent page of a maths paper, or creating true to fact explainers.

They can't.

reply
yorwba
3 hours ago
[-]
Giving it numerical data in an LLM-generated prompt doesn't seem to help much: https://imgur.com/a/KFhczOd

It included the table verbatim and even managed to hallucinate a reasonable heading for it, but then the graph doesn't even manage to align the data points with the time axis, leading to an unfortunate collision in the middle.

I guess you should use a traditional graphing library for your presentation slides for now.

reply
spwa4
3 hours ago
[-]
... but if you want an accurate graph, why not ask the LLM model to put the data points into a graphing library?
reply
Smaily
2 hours ago
[-]
why I cant submitte new posts here ? My account since 2016
reply