Hi,

I haven't uploaded it yet (FTP was down when I finished) although I hope I
will have time to tonight. I will send the link when it is up.

I won't give too many trade secrets away but this is roughly how I do it:

- Get URL for story
- Replace "/hi/" in URL with "/low/"
- Download story HTML
- You may have noticed that the printable version is split into sections by
a <hr>, simple run an array split with <br> as the delimiter (explode() in
PHP I think)
- Then you can run a regular expression to get images out.

I then put them in a MySQL db with a fulltext index so they could be
searchable.

Hope it helps.

David

-----Original Message-----
From: [EMAIL PROTECTED]
[mailto:[EMAIL PROTECTED] On Behalf Of Taha Taha
Sent: 01 June 2005 3:05
To: [email protected]
Subject: RE: [backstage] Content Structure

I know what u mean , for sure the printable version is simpler .
I was hoping to find a 'HowTo book ' of BBC structure.
Where is ur program posts its result anywhere on the web ?

-----Original Message-----
From: [EMAIL PROTECTED] [mailto:[EMAIL PROTECTED]
Sent: May 31, 2005 3:59 AM
To: [email protected]
Subject: Re: [backstage] Content Structure

Hi,

Weird enough, I did this last week!

I too parsed the html of the news item to get the pictures however a little
word of advice: it's ALOT easier to do so when viewing the low detail (i.e.
text
only) version ;)

David

On Mon, 30 May 2005 13:09:12 -0400, "Taha Taha" wrote Hi all, 

�I'm working on a picture outlook for the BBC content , it will display all
the images posted on BBC-with some exceptions-, . I was wondering if there
is any guideline about the structure of the content. 

I read the XML feeds, but when I go into the news items am using my own
observation of how to parse the content (using the internal comments of the
HTML code mostly).Is there a smarter way of doing this or do I have to play
more with the jigsaw?

�

Thanx everyone 

�

----------------------------------------
Scanned by Emailfiltering.co.uk












Reply via email to